acceptodds
Under review as a conference paper at ICLR 2027

MUTUI-BENCH: BENCHMARKING MULTI-TURN IMAGE GENERATION AND EDITING FOR VISUAL INTERFACES

Abstract

A visual interface editor must change the intended element, preserve relevant context, and interpret user intent across turns. We introduce , a compact image-output benchmark with three tracks: Design for stylistic and structural editing, Assist for sequential mouse guidance, and Simulate for rendering browsing-action consequences. The corpus contains 76 tasks and 274 turns: 255 design edits across seven domains and twelve edit types, and 19 interaction turns across six websites. An audit of two archived image-generation systems separates image quality from instruction correctness, removes repeated judgments, and reports matched comparisons with task-level uncertainty. On Design, mean image quality exceeds 8.3/10 while instruction correctness remains below 6.1/10; icon replacement is particularly weak despite polished outputs. The interaction tracks expose complementary demands on pointing and state rendering, although their small size and reference-conditioned evaluation limit generalization. provides an inspectable testbed for studying what should change, what should persist, and what should happen next in visual UI editing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.