Early Context Starvation in Conversational LLMs: Causal Evidence from Attention Interventions
Abstract
Large language models degrade over long conversations, but standard evaluation offers no mechanistic account of why. On human multi-turn dialogue rather than synthetic probe tasks, we show that phase-dependent degradation persists across three generations of two proprietary model families despite real capability gains. On two open-weight models, Gemma 4 31B and Qwen3.6-27B, we then measure and intervene on attention directly. In Gemma, attention to a conversation's earliest content collapses as the dialogue grows, because its sliding-window layers lose sight of early tokens and only its global layers still reach them. A dose-response intervention that biases attention toward or away from early context shows this attention is causal for late-turn accuracy. The effect is position-specific: the same intervention on middle-context keys reverses direction, and on late-context keys produces a sharp threshold collapse. Layer-band interventions localize the early-context effect to Gemma's deep layers. On Qwen3.6 the intervention moves attention just as much but leaves behavior unchanged, because its shallow and deep layers push in opposite directions and cancel. Whether attention to early context matters is therefore a property of the architecture, which motivates phase-resolved evaluation and architecture-aware accounts of multi-turn failure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.