acceptodds
Under review as a conference paper at ICLR 2027

Do We Need Handcrafted Corruptions? Context-Preserving Activation Perturbation in Representation Space

Abstract

Activation patching is widely used in mechanistic interpretability to identify causally important model components. Yet its reliability depends on constructing corrupted inputs that alter task-relevant information while preserving the surrounding context. Designing such inputs is difficult in complex tasks, especially when answer-related evidence is dispersed and related mentions must be edited consistently. To address this challenge, we revisit activation patching in representation space rather than input text and introduce Context-Preserving Activation Perturbation (CPAP), which automatically constructs activation perturbations without using handcrafted corrupted inputs. CPAP maximizes the change in an upstream module's activation while constraining changes in selected downstream residual representations, using their stability as a practical proxy for preserving the surrounding computational context. Our experiments yield two main findings. First, across three models and four benchmarks, CPAP closely matches the component rankings from corrupted-input patching, with consistently higher rank agreement than zero and learned constant ablations. For GPT-2 circuit localization, it achieves the highest AUROC among automatic methods on all three tasks (0.968–1.000). Second, CPAP easily extends to explaining model behavior across complex tasks and larger models (27B). In long-context question answering and mathematical reasoning, it automatically constructs activation perturbations and identifies key components without manually designed corrupted inputs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.