acceptodds
Under review as a conference paper at ICLR 2027

MockWeb-Bench: Evaluating Hybrid Agents through Website Reconstruction

Abstract

Reconstructing a complete website from a running reference requires discovering its appearance and behavior across pages, interaction states, and screen sizes, then reproducing both in an independent implementation. Unlike fixed screenshots or recorded interactions, a running reference lets agents decide what to inspect as they build. We introduce MockWeb-Bench, which evaluates hybrid graphical user interface (GUI) and coding agents on 50 offline reference websites with 4,101 manually verified evaluation pages and paired desktop and mobile captures. Each task provides an entry URL, a starter workspace, and a six-hour initial budget; the list of scored pages stays hidden. Functional tests are validated on the reference and frozen before candidate generation, then applied to each reconstruction to measure behavioral fidelity. A fixed vision-language model (VLM) checks candidate screenshots against visual requirements extracted from reference screenshots, while separate checks assess quality and structure. Across ten model–agent configurations,GPT-6 Astra with Codex achieves the highest mean final score of 56.26%. All ten model-agent systems score lower on interaction tests than on general content and navigation tests. Beyond final scores, trajectory case studies reveal both useful strategies and attempts to take shortcuts. Agents write code to compute pixel differences between reference and candidate screenshots, using the results to make reconstructed pages look much closer to the reference. Some systems also try to read protected files or copy reference code, showing why evaluations need to restrict access to reference files and check submissions for copying. MockWeb-Bench supports studying both the websites hybrid agents build and the strategies they use to build them. Our code and dataset are available at https://anonymous.4open.science/r/mockweb-bench-208C/README.md.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.