acceptodds
Under review as a conference paper at ICLR 2027

Do Small Errors Snowball?Understanding and Controlling Error Propagation in Language Agents

Abstract

Small observation changes can propagate through a language agent’s feedback loop, changing its actions, environment state, and later inputs. Terminal success conflates initial sensitivity with subsequent amplification, while one-step agreement cannot measure recovery. We introduce an affine Wasserstein profile in a policy- independent task-semantic geometry that separates leakage on equivalent states from amplification on separated states, with composition and first-exit bounds linking local dynamics to outcomes. Coupled-Rollout Contraction Optimization (CRCO) then trains marginal-preserving coupled continuations, penalizing suc- cessor separation only beyond a state-conditioned allowance. We prove that this correspondence distinguishes pair laws that absolute and constant-threshold penal- ties cannot, and that the optimized hinge bounds task-score sensitivity with a finite-sample guarantee. Across nine backbones and eight interactive settings, CRCO has the highest perturbed-score point estimate among six matched objec- tives on all nine backbones. It exceeds matched absolute propagation by 2.6/3.3 points on two instrumented anchors, retaining a 2.6-point mean gain with an observation-trained metric without latent predicates. Separately analyzed reference- prefix and reset-to-terminal protocols yield equal-cell mean gains of 3.2 and 2.5 points. Threshold reassignment, recovery interventions, and five-step checkpoint selection corroborate the mechanism; deployment remains one ordinary rollout.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.