GUI-GenBench: Evaluating Image Generation Models as Interactive GUI Environments
Abstract
Recent advancements in image generation models enable the prediction of future Graphical User Interface (GUI) states based on user instructions. However, existing benchmarks primarily focus on general-domain visual fidelity, leaving evaluation of state transitions and temporal coherence in GUI-specific contexts underexplored. To address this gap, we introduce GUI-GenBench, a benchmark for evaluating dynamic interaction and temporal coherence in GUI generation. GUI-GenBench comprises 700 carefully curated samples spanning five task suites, covering both single-step interactions and multi-step trajectories across real-world and fictional scenarios, as well as grounding point localization. To support systematic evaluation, we propose GUI-Score, a five-dimensional metric that assesses Goal Achievement, Interaction Logic, Content Consistency, UI Plausibility, and Visual Quality. Extensive evaluation indicates that current models perform well on single-step transitions but struggle with temporal coherence and spatial grounding over longer interaction sequences. Moreover, our findings identify icon interpretation, text rendering, and localization precision as key bottlenecks, and suggest promising directions for future research toward high-fidelity generative GUI environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.