acceptodds
Under review as a conference paper at ICLR 2027

PerceptArena: Benchmarking Perception as a Decision in Web Agents

Abstract

Computer-use benchmarks have proven sharp instruments for measuring an agent's execution ability: given an observation, can it select the right actions? Yet nearly all of them share a single harness design: the environment is rendered and fed to the agent at every step, so the agent itself never decides what to observe. We introduce PerceptArena, which returns that control over perception to the agent itself: how much to observe, at what fidelity, and when to stop observing. PerceptArena is a controllable, self-hosted web environment comprising nine scenarios and a cross-scenario layer with 652 tasks in total; each feasible task places its decisive evidence at particular locations and information carriers, making perception allocation part of the task rather than a fixed input. We evaluate 18 vision-language models spanning frontier and mid-tier APIs under both the mainstream Harness-Supplied Observation protocol (HSO) and our proposed Agent-Controlled Perception protocol (ACP). Seventeen of the eighteen models attain a higher cost-success AUC under ACP, and sixteen spend less metered tokens; the median model spends 59% of its HSO metered tokens, average success is essentially unchanged, moving 2.1 points field-wide and 0.9 points with the two degenerate baselines excluded, and evidence recall across the field rises from 60.0% to 64.7%. Nevertheless, ACP also exposes a shared weakness of frontier models in perception self-management: models struggle to target the decisive evidence, track what they already hold, and stop once the evidence suffices. These observations remain hidden under the traditional HSO protocol. PerceptArena thereby makes perception allocation a measured behavior, and turns what to look at, what to keep, and when to stop into explicit, trainable targets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.