Lost in the Context: Why Transformers Fail to Revisit the Right State
Abstract
Transformers can attend to their entire preceding context, yet they often fail on long-horizon reasoning tasks that require keeping and revisiting intermediate computational states. We distinguish two possible bottlenecks: state persistence (does a state written earlier still influence computation at later tokens and layers?) and referential persistence (can the model later retrieve it from among competing states?). Taking a dynamical-systems view in which token-wise transformations act as reaction terms and attention acts as a non-local transport and retrieval operator, we treat recall as depending on how much of a state reaches the query, how the query reads it out, and how it is selected among competitors. We measure the first as the linear response of the residual stream to a perturbation of one token's embedding; as earlier Green's-function analyses found, it decays roughly as a power law, with an exponent that falls with depth. Post-training can strengthen this long-range influence: in the Qwen pair, RL slows its decay on a reasoning-style probe through a broad shift of attention toward long range in middle and late layers (exponent change −0.021, replicated at −0.014, robust to a constant floor in the fit and about 75 times the change from random updates of the same size), but not on raw text. Stronger influence, however, does not mean better recall. An OLMo DPO step raises the state's influence 1,000 tokens away by 73% and the answer's sensitivity to it by 24%, yet lowers recall from 97.5% to 85% (15 prompts lost, none gained), and across prompts we detect no reliable link between the decay exponent and which recalls fail. Head ablation makes the dissociation causal: removing the heads that carry the most long-range influence steepens decay but leaves recall intact, whereas other head sets, without raising perplexity, abolish recall when the value must be selected among competing assignments. Long-range influence is therefore not a proxy for memory: methods meant to help models revisit earlier states should be judged by retrieval and readout, not by how far influence travels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.