acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Fitted Answer: Causal Dissection and Downstream Consequences of Learned Activation Interventions

Abstract

Activation interventions are a widely used method to associate language model (LM) representations with interpretable variables. However, success on the objective used to fit an intervention neither establishes that the intervention produced the intended downstream consequences, nor does it identify which internal components caused the behavioural changes. We investigate these questions by constructing a Theory of Mind (ToM) belief-tracking dataset comprised of short narratives and associated questions about agents’ beliefs and their consequences. Rank-16 activation patches are trained under the objectives of 1) transferring the location answer from a source narrative and 2) implementing a consistent remapping of locations. Evaluations of the activation patches are carried out by measuring the differences in matched target rules, held-fixed downstream consequence tests, and with unfitted principal component analysis (PCA) controls. Answer accuracy is shown to be insufficient to establish a coherent intervention. On the consequence test set, a PCA intervention followed by a fixed output relabelling achieves 100% accuracy on the target question under the alternative mapping, yet fails to answer all follow-up questions correctly for any narrative. By contrast, holding each patch constant across questions and subsequent observations, we find that the learned remapping led to the LMs behaving as if the underlying events had been rewritten, showing high agreement with follow on consequences, including those not trained for. Nevertheless, comparisons with PCA controls reveal no additional accuracy benefit from fitting for the intended source-answer transfer objective on these consequence-based evaluations. To test the causal contribution of attention keys to the behavioural difference between learned and PCA interventions, we interchange the attention keys induced by the two interventions. Queries and values evolve within the recipient. Exchanging the PCA-induced attention keys with the learned patch attention keys led to increased agreement with the trained patches’ outputs by up to 75.4 percentage points, averaged over multiple seeds, whereas the reciprocal exchanges favoured PCA outputs. On paired colour-lookup questions, performance also exceeds an exact upper bound achievable by any single table-independent transformation of the PCA colour answer. Together, these results establish the tested downstream scope of learned remapping and provide causal evidence that attention keys induced by the learned patch contribute to the behavioural differences between learned and PCA interventions under the tested multilayer intervention setting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.