acceptodds
Under review as a conference paper at ICLR 2027

WebChallenger: A Reliable and Efficient Generalist Web Agent

Abstract

Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference cost is prohibitive for the repetitive tasks where such agents would be most useful. We argue this gap stems not from insufficient model capability but from agent architectures that fail to replicate three human cognitive advantages: selective attention to relevant page regions, persistent memory of website structure, and procedural fluency with common interaction patterns. We introduce WebChallenger, a web agent framework that addresses each gap through architecture design rather than model scale, built around PageMem: a structured page representation deterministically constructed from the DOM that exposes each page as a hierarchy of semantic sections with short summaries. On this substrate we build three mechanisms that mirror these advantages: an observation pipeline that reads only task-relevant sections, a memory built by a one-time offline crawl of each website, and workflows that execute multi-step interactions as single actions. Because all three operate over PageMem, the framework generalizes across websites without site-specific adapters. Using off-the-shelf open-weight models without fine-tuning, our system achieves 56.3% on WebArena, 48.7% on VisualWebArena, 51.0% on Online-Mind2Web, and 70.9% on WorkArena, approaching frontier proprietary systems at a fraction of the cost. With the model held fixed in each comparison, WebChallenger scores 39 points higher on WebArena-lite than the standard BrowserGym harness and 14 points higher than the top-ranked WebArena harness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.