acceptodds
Under review as a conference paper at ICLR 2027

INTERVENTION IS NOT IDENTIFICATION: ACTIVATION EDITS CAN PRODUCE THE PREDICTED EFFECT FOR THE WRONG REASON

Abstract

An activation intervention can be exact, task-shaped, and sham-specific without identifying the mechanism that produced its effect. We show this in selective rule updating, where a current binding B and applicability variable S determine the required action A= B⊕S. Qwen2.5-14B-Instruct executes the operator perfectly when both operands are explicit (400/400). A free-reasoning scaffold raises accuracy to 89.6%, but it changes the computational trajectory; direct full-context accuracy is 46.2%. A bitwise-verified layer-28 replacement of the decoded B coordinate moves a supervised layer-35 readout of Aby +0.0288 and reverses across applicability conditions as the task predicts. The effect survives identity, zero-edit, and equal-norm orthogonal controls. It also persists when the measured A coordinate is held fixed at the edit site (paired held-minus-free difference +0.0011, 95% CI [−0.0021,+0.0043]). These checks establish a real, structured internal effect. They do not identify its mechanism. Under the pre-specified decomposition, contraction toward the downstream readout mean accounts for 74.3% of the interaction, while the baseline-independent remainder is unresolved. On cases where contraction and task-directed correction predict opposite signs, the displacement is near zero. Recipient dependence exceeds the task-directed bound by 4.58×and 4.18×. A structured task-irrelevant variable is more decodable than the target, simple raw-vector carry-through explains essentially none of the downstream interaction, and no preferred answer changes. Across 1,000 matched synthetic draws per condition, the directed estimator is approximately unbiased with 94.2% empirical coverage and separates injected directed effects from pure contraction under the specified generating process. The result separates a causal intervention effect from the mechanism used to explain it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.