MUTUI-BENCH: BENCHMARKING MULTI-TURN IMAGE GENERATION AND EDITING FOR VISUAL INTERFACES
Abstract
A visual interface editor must change the intended element, preserve relevant context, and interpret user intent across turns. We introduce , a compact image-output benchmark with three tracks: Design for stylistic and structural editing, Assist for sequential mouse guidance, and Simulate for rendering browsing-action consequences. The corpus contains 76 tasks and 274 turns: 255 design edits across seven domains and twelve edit types, and 19 interaction turns across six websites. An audit of two archived image-generation systems separates image quality from instruction correctness, removes repeated judgments, and reports matched comparisons with task-level uncertainty. On Design, mean image quality exceeds 8.3/10 while instruction correctness remains below 6.1/10; icon replacement is particularly weak despite polished outputs. The interaction tracks expose complementary demands on pointing and state rendering, although their small size and reference-conditioned evaluation limit generalization. provides an inspectable testbed for studying what should change, what should persist, and what should happen next in visual UI editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.