IUX-GenBench: Benchmarking Interactive User Experience for Multi-Turn Image Generation
Abstract
Existing benchmarks for image generation and editing models have made considerable progress in evaluating static output quality across diverse tasks. Nevertheless, such evaluations tend to concentrate on isolated capabilities, thereby neglecting the dynamic interaction requirements and multifaceted scenarios that users encounter in practice. We therefore introduce **IUX-GenBench**, a **bench**mark for **i**nteractive **u**ser e**x**perience in multi-turn image **gen**eration and editing tasks. Our main contributions include: (i) constructing a hybrid dataset of UGC, PGC, and LLM-synthesized augmented data with a fine-grained taxonomy of capabilities and scenarios; (ii) introducing a multi-model collaborative simulator to replace static evaluations with dynamic assessment across simulated interaction loops; (iii) designing a dual-layer assessment system that jointly measures turn-level quality and session-level coherence; and (iv) proposing automated attribution analysis and structured diagnostic reporting as a novel evaluation practice that transcends scalar scoring to deliver interpretable, actionable failure diagnostics. Extensive experiments show that most models perform poorly on IUX-GenBench, with significant session-level degradation. Attribution analysis further reveals that unintended background alterations and ignored local constraints account for most session-level failures, highlighting persistent deficiencies in sustaining user intent and contextual coherence. By exposing these limitations, IUX-GenBench provides a demanding testbed for generative AI development.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.