acceptodds
Under review as a conference paper at ICLR 2027

Affect on Demand: Causal Retrieval of Character Affect in a Language Model

Abstract

Language models can use a character's earlier success or failure after several neutral sentences, but behavioral success leaves open a mechanistic question: is the relevant affect-linked state carried continuously, or retrieved from the earlier episode when it becomes useful? We distinguish these possibilities with a prospectively frozen four-stage intervention pipeline in Qwen2.5-1.5B-Instruct. A held-out representational confirmation first finds both persistent maintenance and cue-dependent reinstatement (maintenance = 1.648 diagnostic-TRAIN SD; retrieval-dependent index = 0.289; both Holm-adjusted ), showing that decoding alone does not select a single mechanism. We then intervene on source-span key/value states. Opposite-affect source patching at a development-selected layer produces positive cue-specific causal restoration (CRE = 0.0397, 95% CI [0.0371, 0.0424]) and strong source-over-bridge specificity (0.0593, [0.0566, 0.0620]). Head-level development localizes a two-head K/V pathway (L12H0, L13H0); protected confirmation supports perturbation-relative necessity, matched-head specificity, and restoration rescue. Finally, an update experiment shows selective retrieval after irrelevant suppression (, [0.127, 0.245]), while a contradictory update does not produce the prospectively frozen outdating signature. Across levels of analysis, the evidence supports a hybrid mechanism: affect-linked information remains decodable across discourse, yet later use depends causally on retrieving source-bound states through a compact pathway rather than simply reading out a continuously overwritten scalar state.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.