acceptodds
Under review as a conference paper at ICLR 2027

ViRE-Bench: Benchmarking Repository-Level Frontend Code Editing from Visual Feedback

Abstract

Frontend editing requests often concern visible content, appearance, and spatial relationships. Annotated screenshots ground these requests directly in the rendered interface, allowing users to communicate changes to coding agents without identifying source files or components. Agents must then infer the intended edit, locate the relevant repository code, and implement the change while preserving unrelated content and appearance. We introduce ViRE-Bench, a benchmark of 1,000 repository-level frontend code-editing tasks drawn from 202 real-world HTML, React, Vue, and Angular repositories. The benchmark spans eight edit operations, covers both instance-level and shared-component edits that propagate across pages, and evaluates submission validity and intent satisfaction alongside rendered measures of target fidelity and preservation. Among six models evaluated under a common single-turn protocol, the best-performing model achieves an overall benchmark score of 67.29 and an intent satisfaction rate of 70.2%. Component movement and resizing remain challenging for all six models. Controlled diagnostic experiments show that explicit editing intent yields larger gains than source-location information alone across all three tested models, with two models reaching approximately 95% intent satisfaction under explicit intent. These findings indicate that deriving an actionable editing specification from visual feedback is a primary source of difficulty, while the benefits of source-location assistance vary across models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.