acceptodds
Under review as a conference paper at ICLR 2027

Long-Horizon Web Agents Need a State Beside the Trajectory

Abstract

LLM agents struggle on long-horizon tasks for two reasons. First, the policy fails to attend to evidence that is scattered across a long trajectory yet needed for the next decision. It tends to commit to actions without grounding them in the current evidence. Second, when the current environment differs from training in ways the trajectory does not reveal, the policy falls back on training-time priors rather than checking the current environment. We propose Cast (Coded Agent with STate), an agent loop that maintains a small, structured state alongside the trajectory and exposes it to the policy at decision time. This state collects evidence relevant to the next decision as the trajectory unfolds, so the policy reads it from one place rather than recovering it from the trajectory. Environment-side fields are either filled from the current observation or marked as unknown, so an unresolved fact is visible to the policy before it acts. Deterministic code updates the state after every step with no language-model call, and a fixed contract ties a few state values to behavioral constraints. On WebArena, Cast improves success rate across backbones without any offline exploration corpus. Under targeted environment perturbations spanning latent shifts, execution noise, and structural changes, Cast outperforms the baseline. Our code is available at: https://anonymous.4open.science/r/cast-5DB6/README.md.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.