acceptodds
Under review as a conference paper at ICLR 2027

ProductWebBench: Evaluating Website Continuity Beyond Task Completion

Abstract

Frontend coding agents are usually graded on one question: did the requested target get built. Real engineering changes carry a second obligation, because the rest of the site has to keep working. We call that second obligation *website continuity*, and introduce ProductWebBench, a benchmark that scores both on living frontend repositories: 320 Change tasks over existing code and 80 long-horizon Build tasks. Every task freezes a repository snapshot, staged requirements, browser evidence and deterministic checks before evaluation, and enters only once a reference passes every check and plausible-wrong ones fail. A simulated user reveals the requirements one stage at a time, and every check cleared earlier stays armed for the rest of the run. Across 8,737 runs from 13 models the two come apart: of the Change runs passing every requested-work check, **28.2%** still fail a continuity check — a blank route, horizontal overflow, console errors — and **10.8%** push a browser state they had already passed back to failing. Because the same checks also run on the untouched repository, each regression carries exact provenance, not a post-hoc label: on the 586 runs where both readings exist, **18.3%** are caught breaking a verified state. Interaction is a further blind spot, failing in **29.5%** of runs whose static DOM already passes. ProductWebBench therefore scores a patch against the living system it lands in, not the target alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.