acceptodds
Under review as a conference paper at ICLR 2027

Context Rot in Agent Trajectories: Evolving-State Maintenance Is the Hidden Bottleneck of Long-Context Agents

Abstract

Long-context language models are increasingly deployed as agents whose context is not a static document but a self-generated multi-turn trajectory of (thought, action, observation) steps. We ask whether early key information is faithfully used as this trajectory grows. We introduce a controlled probe suite that holds task difficulty fixed while varying only trajectory length K in 1,...,128, across four task families that separate passive recall (key–value lookup, early constraints, multi-hop) from active evolving-state maintenance. Evaluating nine frontier models, we find a sharp dissociation: passive recall is largely robust (constraint and multi-hop stay at ceiling; KV recall degrades mildly), whereas evolving-state maintenance collapses catastrophically for every model once K >= 16 (early-information utilization drops from  1.0 to  0.0; all decay slopes lambda < 0). We further isolate an agent-specific corruption source: a model's own erroneous intermediate conclusion written into the trajectory (self-conflict) degrades performance significantly more than neutral filler of equal length (Delta = -0.150, 95% CI [-0.199,-0.102]); a controlled comparison shows this is driven by the misleading content rather than self-authorship per se, except for the strongest reasoner. Moreover 47.5% of these failures land exactly on the "snowball" value the model would obtain by faithfully continuing from its own error—an effect that worsens with reasoning strength. A two-phase mechanism emerges: shallow-K failures are dominated by self-generated-error snowballing (which a from-scratch recompute prompt partially repairs for weaker models, e.g. +0.333 for Claude Haiku), while deep-K failures are dominated by forgetting (which no training-free prompt repairs; a mechanism-targeted recompute recovers +0.211, 95% CI [+0.139,+0.289], in the shallow snowball regime but +0.007, not significant, at the deep cliff). We corroborate both the cliff and the self-error effect on genuine multi-turn rollouts in which each model self-generates its own trajectory: the collapse to EUR 0 reproduces, and rollouts containing a self-authored intermediate error fail 6.1x more often. Together these establish evolving-state maintenance as a stubborn, agent-specific long-context bottleneck and provide a first partial solution scoped to where it works.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.