acceptodds
Under review as a conference paper at ICLR 2027

When Activation Interventions Fail to Identify Component Rankings

Abstract

Mechanistic interpretability often treats a successful activation intervention as evidence that a component's causal role has been identified. We show this inference can fail for the reported ranking itself: for a fixed finite intervention family, a second decomposition can match every observed outcome exactly yet reverse the ordering of two components on a held-out contrast, while leaving every residual-stream activation and output unchanged. The ambiguity is governed by coverage, the portion of the reported contrast outside the span of the probed directions, measurable on the original model before any alternative is built. On GPT-2 Small, a construction reverses the IOI name-mover ordering that single-position interchange measures on all 24 held-out pairs while agreeing with a 512-prompt probe family, and generalizes to unseen templates. On Gemma-2 2B, an edited same-architecture JumpReLU SAE matches reconstruction and every probe outcome where the intervened feature is inactive, while reversing the feature-pair ordering on most active prompts. Coverage separates evidence styles: single-position families can leave the ordering steerable until probe count nears the model width (768 on IOI here), while position-resolved families and the per-context trace of a discovery sweep drive the uncovered portion to or near zero at this width. Successful interventions alone do not establish identification; whether the evidence covers the claim does.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.