acceptodds
Under review as a conference paper at ICLR 2027

HORIZON: Evidence-Grounded Task Memory, Interface Translation, and Experience for Long-Horizon Web Agents

Abstract

Long-horizon web tasks require agents to connect user requirements to the effects of actions across changing interfaces. An executed action can leave the requested outcome unresolved: a click may activate the wrong control, a value may remain unsaved, or a returned answer may lack supporting observations. We introduce HORIZON, a framework that makes these connections explicit through evidence-grounded task state. Task Completion Memory tracks requirements, object bindings, progress, and supporting observations. A Document Object Model (DOM) Translation Layer links structured interface views and semantic operations to their source elements. State-conditioned experience retrieves prior task records as advice and provides a formulation for contrasting successful and unsuccessful interactions while keeping historical advice separate from current-task evidence. On WebArena, HORIZON raises gpt-5.6-luna's overall success rate from 15.4% to 40.3%, a gain of 24.9 percentage points. With Qwen3.8-27B, it achieves 60.1% across all 812 tasks, including 66.7% on cross-site tasks, under a protocol with up to two attempts on four single-site domains. These results compare complete agent configurations. Selected trajectories illustrate source-linked targeting, requirement-directed readback, and archived-route adaptation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.