IterUI: Learning DOM-Grounded Local Visual Code Editing for UI-to-Code
Abstract
Vision-language models (VLMs) can generate executable webpages from UI screenshots, yet correcting their remaining visual discrepancies remains challenging. Existing full-page refinement treats each correction as another page-generation step, leaving the correspondence between a local visual error and its responsible code implicit and potentially disturbing already-correct regions. We instead formulate UI refinement as local visual code editing, where each action pairs a grounded DOM component with an executable code update whose effect is evaluated after full-page rendering. Based on this formulation, we propose IterUI, a closed-loop framework that detects discrepancies between a reference screenshot and the current rendering, grounds them to DOM components, and generates targeted patches from localized visual evidence, component code, and ancestor layout context. IterUI masks non-target pixels while preserving the original image canvas, retaining the position and scale information needed for visual correction. Each patch is reintegrated into the complete page and accepted only after render-based verification. Because the framework requires only editable source code, browser renders, and the rendered DOM, it can also refine outputs from existing VLMs without updating their parameters. We further train IterUI-Coder-8B on refinement trajectories and optimize its code-generating actions using Group-Scored Relative Policy Optimization (GSRPO), which jointly evaluates groups of rendered candidates to construct comparative visual rewards. Across four UI-to-code benchmarks, using a test-time judge distinct from the training reward model, IterUI-Coder-8B achieves the highest VLM scores among the evaluated open-source models. On UIHTML, IterUI further raises its VLM score from 66.80 to 69.12.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.