WebProgramBench: Can Coding Agents Reconstruct Live Applications from Scratch via Black-Box Interaction?
Abstract
Coding agents have recently achieved substantial progress in fixing bugs and implementing features within existing repositories, yet building complete real-world applications from scratch remains an open challenge. Existing software-development benchmarks assess only partial aspects of this capability through requirement-driven generation and screenshot-based interface cloning. Their evaluations emphasize visual similarity, selected functional checks, or model-based judgments, leaving room for shortcut solutions that pass evaluation. To address these limitations, we introduce WebProgramBench, the first execution-based benchmark for evaluating whether LLMs can reconstruct real-world live applications from scratch through black-box interaction. WebProgramBench requires agents to actively explore a live reference application and reconstruct its behavior in a deployable implementation, without access to its source code. To evaluate task performance, we introduce a novel cross-layer protocol that jointly verifies interface, service, and persistence behavior, preventing frontend-only mocks from passing. The benchmark comprises 20 real-world applications spanning four domains, with reference implementations totaling nearly one million lines of code. To scale this evaluation, our fully automated pipeline generates 1,254 verified executable tests with 20,851 assertions. Our experiments across five leading LLMs show that frontier LLMs remain far from complete real-world application reconstruction. Even the best-performing model achieves near-complete reconstruction on only three applications, only one of which is fully reconstructed. Together, these results establish WebProgramBench as a rigorous benchmark for measuring LLM capabilities in autonomous, closed-loop software engineering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.