HORIZON: Evidence-Grounded Task Memory, Interface Translation, and Experience for Long-Horizon Web Agents
Abstract
Long-horizon web tasks require agents to connect user requirements to the effects of actions across changing interfaces. An executed action can leave the requested outcome unresolved: a click may activate the wrong control, a value may remain unsaved, or a returned answer may lack supporting observations. We introduce HORIZON, a framework that makes these connections explicit through evidence-grounded task state. Task Completion Memory tracks requirements, object bindings, progress, and supporting observations. A Document Object Model (DOM) Translation Layer links structured interface views and semantic operations to their source elements. State-conditioned experience retrieves prior task records as advice and provides a formulation for contrasting successful and unsuccessful interactions while keeping historical advice separate from current-task evidence. On WebArena, HORIZON raises gpt-5.6-luna's overall success rate from 15.4% to 40.3%, a gain of 24.9 percentage points. With Qwen3.8-27B, it achieves 60.1% across all 812 tasks, including 66.7% on cross-site tasks, under a protocol with up to two attempts on four single-site domains. These results compare complete agent configurations. Selected trajectories illustrate source-linked targeting, requirement-directed readback, and archived-route adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.