When Activation Interventions Fail to Identify Component Rankings
Abstract
Mechanistic interpretability often treats a successful activation intervention as evidence that a component's causal role has been identified. We show this inference can fail for the reported ranking itself: for a fixed finite intervention family, a second decomposition can match every observed outcome exactly yet reverse the ordering of two components on a held-out contrast, while leaving every residual-stream activation and output unchanged. The ambiguity is governed by coverage, the portion of the reported contrast outside the span of the probed directions, measurable on the original model before any alternative is built. On GPT-2 Small, a construction reverses the IOI name-mover ordering that single-position interchange measures on all 24 held-out pairs while agreeing with a 512-prompt probe family, and generalizes to unseen templates. On Gemma-2 2B, an edited same-architecture JumpReLU SAE matches reconstruction and every probe outcome where the intervened feature is inactive, while reversing the feature-pair ordering on most active prompts. Coverage separates evidence styles: single-position families can leave the ordering steerable until probe count nears the model width (768 on IOI here), while position-resolved families and the per-context trace of a discovery sweep drive the uncovered portion to or near zero at this width. Successful interventions alone do not establish identification; whether the evidence covers the claim does.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.