WebMockGym: Benchmarking Web Reconstruction with Autonomous Information Gathering
Abstract
Autonomous software development requires agents to determine what they need to know, acquire and organize relevant information, and use it to produce reliable implementations. Evaluating end-to-end development calls for tasks that leave exploration and implementation decisions to agents while providing clear criteria for judging the final result. Existing web-development benchmarks often supply screenshots or interaction recordings, offering limited evidence about development when agents must gather reference information themselves. We introduce WebMockGym, a benchmark of 50 websites and 50 web apps for studying autonomous development through website reconstruction. Agents explore a frozen reference website and build a separately runnable reconstruction of its entry page and captured first-hop pages. The reference provides an observable target, while agents decide what to inspect and how to implement and check their work. Fixed, hidden tests assess content, images, visual appearance, and cross-page navigation. The same reference-defined requirements apply to every reconstruction of a task, including requirements an agent may overlook. Experiments with six coding-agent systems reveal frequent navigation failures, and no system passes every applicable test on any scored task in the website track. Trajectory analysis associates earlier writing and more frequent inspection of the reconstruction followed by further writing with higher Web scores. WebMockGym provides a concrete setting for studying autonomous development through the quality of independently runnable reconstructions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.