acceptodds
Under review as a conference paper at ICLR 2027

CONTROLS ARE INTERVENTIONS: WHAT COMPLEMENT ABLATIONS ERASE, AND HOW TO AUDIT THEM

Abstract

Circuit analyses often replace everything outside a circuit with mean, zero, or resampled activations, then knock out parts of it. This complement ablation is treated as a control but is an intervention, erasing each input’s own activations and the replaced components’ reactions to knockouts. We audit it against the full model on knockout effects and the head choices they induce. In three published GPT-2 circuits, for indirect object identification (IOI), acronyms, and greater-than, mean-ablating the non-circuit heads keeps clean accuracy within five points. Yet on every fresh input set, its per-input choice of the most damaging head pair misses more full-model effect than our tolerance allows. Freezing these heads at each input’s clean activations keeps clean outputs exact and blocks only their reactions; it removes most of this error in IOI, about half elsewhere. Unfreezing the 48 non-circuit heads that attribution patching ranks highest passes in IOI and greater-than, not acronyms. In the original completeness test of the published IOI circuit, knockout sets found by greedy search have large gaps under mean ablation but small ones under freezing; a frozen-complement search finds the reverse. In IOI, freezing attention to the first token, the attention sink, removes over a third of the self-repair after zeroing the name-mover heads. Auditing the decisions a complement ablation supports, not its clean behavior, makes its faithfulness testable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.