On the Stability of Causal Attribution Across Language Model Continuations
Abstract
An internal state’s measured influence on a language model’s behavior may depend on the continuation that follows it. We test this dependence using error acknowledgment in mathematical reasoning. Using matched clean–error traces, we replace regional key–value cache states after prefill across eight checkpoints in thinking and non-thinking modes. Effects on acknowledgment concentrate in the localized error span and post-error states, with thinking generally increasing relative error-span sensitivity. Regional effects on an early error-related projection differ systematically from those on acknowledgment: error-span replacement has a larger relative effect on acknowledgment, whereas post-error replacement shows the reverse pattern. Across six checkpoints, replaying natural thinking prefixes partially reproduces the thinking-associated increase in error-span sensitivity under non-thinking generation, while controlled artificial prefixes show that selected continuation content changes the balance between error-span and post-error sensitivity. These prefix interventions hold the checkpoint, original prompt prefill, and target decoding mode fixed. Together, these findings show that, in this setting, regional causal attributions obtained through cache replacement depend on subsequent continuation content and need not transfer unchanged across continuations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.