LongWebBench: Fine-Grained Evaluation of Agents on Long-Horizon Web Tasks
Abstract
Agents powered by large language models are increasingly asked to complete long-horizon tasks that span hundreds of dependent interactions. Evaluating them calls for tasks that are long by construction, an environment that is controllable yet disrupts the agent as real runs are disrupted, and an evaluation that shows where a run falls short, but existing benchmarks rarely provide all three. To fill this gap, we introduce LongWebBench, a benchmark of long-horizon web data collection on generated, self-hosted websites. Each of its 85 tasks, spanning five difficulty levels, asks the agent to collect a complete dataset, such as every product of an online store with all its attributes, amounting to several hundred to a few thousand items spread over listing, detail, and login-protected pages. During collection, verification gates, expiring sessions, truncated responses, and obfuscated pages disrupt the run. Because the server records what it delivers, LongWebBench compares what a task requires, what the server delivered, and what the agent submitted, and it diagnoses each run at four stages, namely Interface Access, State Control, Data Acquisition, and Result Assembly. We evaluate ten agent systems on the benchmark and find that even the strongest remains far from complete collection and that rankings on the easiest level do not carry over to the hardest. The diagnosis further shows that agents fail mainly by omission, since most of the missing information was never delivered to them rather than misread, and that the backbone model shapes what is collected far more than the agent harness does. A stage-wise evaluation grounded in delivery records thus reveals what a single score hides and where agents need to improve.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.