WebsiteBench: Can AI Agents Rebuild Websites through Browser Exploration?
Abstract
Humans learn how unfamiliar environments work by observing them, trying different actions, and revising their understanding based on what happens. Websites provide a concrete setting for studying this process: they combine interfaces, application logic, and persistent data, while each browser interaction reveals only part of their behavior. Can AI agents decide what to explore and use what they learn to rebuild a website faithfully? We introduce WebsiteBench, which contains 74 expert-built and expert-reviewed website reconstruction tasks and 50,356 hidden tests. In each task, an agent explores a locally hosted website through a browser and builds a runnable offline site that looks and behaves the same way. The agent decides which pages to visit and which interactions to try, and may revisit the reference website while coding and testing. The hidden test suite consists of three groups: Request checks page responses and form submissions; Browser checks what users see and can do; and System checks application requirements. Our experiments show that WebsiteBench is challenging for frontier models: even the top model, Claude Opus 5.5, scores only 22.5%; GPT 5.6 Sol scores 15.6%, and Kimi K3 scores 8.9%. We systematically analyze agent trajectories and identify three failure modes: (1) underexploration, (2) failure to implement observed behavior, and (3) early stopping. We also find that agents given human exploration traces achieve higher reconstruction scores. These results show that agents struggle to learn how an unfamiliar system works through interaction and to use that understanding to rebuild it faithfully.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.