acceptodds
Under review as a conference paper at ICLR 2027

VISTA: Looking Right Is Not Working Right — Benchmarking Coding Agents from Design Handoffs to Full-Stack Web and Mobile Apps

Abstract

Coding agents can now build applications from design mockups, but a screen that looks right is not an application that works. Existing benchmarks either score visual reconstruction without a backend or test full-stack functionality from textual requirements, and few tie scores to the individual components a design requires. We introduce VISTA (VIsual Spec-To-App), an end-to-end benchmark in which coding agents turn multi-page design handoffs (Figma renders, structure, and textual requirements) into runnable full-stack Web and Mobile (Android) applications. VISTA provides 8,487 human-annotated interactive components across 18 applications (126 Web pages and 576 Android screens) and an executable evaluator that locates annotated components in the running application and probes their interactions, yielding localization, behavior, and joint scores traceable to individual design requirements. Across 14 deployed coding-agent systems, the best joint score is 0.553 on Web and 0.378 on Mobile, and 13 of 14 systems score lower on behavior than on localization on Web. Component probes detect unresponsive controls in 90.1% of audited Web deliveries. Development traces show where the process departs from the task: 44.5% of trajectories never open the design screenshots, 73.0% omit the prescribed self-audit, and 8.3% report success after a failed check. VISTA makes the gap between rendering a design and delivering a working application measurable at the level of individual components.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.