acceptodds
Under review as a conference paper at ICLR 2027

WebAppBench: Benchmarking Functionality of Agent-Generated Web Applications

Abstract

Agentic coding systems increasingly generate complete Web applications from natural-language requirements. Evaluating these applications, however, remains challenging: successful deployment and visually plausible interfaces do not establish that users can complete intended workflows across interactions, pages, reloads, and evolving application state. Fixed test scripts are difficult to transfer across open-ended implementations, while generic browser exploration does not by itself provide a requirement-level functional oracle. We introduce WebAppBench, a bilingual, specification-grounded benchmark and evidence-grounded evaluation framework for functional correctness in agent-generated Web applications. WebAppBench contains 100 Chinese and English specifications decomposed into 3,197 typed atomic functional checkpoints. Each checkpoint defines a behavioral claim, an explicit verification condition, and a criticality level, covering navigation, state transitions, value correctness, persistence, perceptual requirements, and runtime integrity. Given a deployed application, a browser-based computer-use agent executes task-relevant workflows without implementation-specific selectors and grounds each verdict in action trajectories, observed state changes, reload outcomes, and runtime signals. WebAppBench records whether each requirement is supported, contradicted, or left with insufficient evidence, making aggregate scores traceable to individual user requirements and their supporting evidence. Experiments with representative LLM and coding-agent configurations enable fine-grained analyses of functional capabilities and failure patterns beyond build status, screenshots, or holistic application-level scores.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.