acceptodds
Under review as a conference paper at ICLR 2027

Intervention-Aware Compression of Neural Networks via Causal Abstraction

Abstract

Model distillation and mechanistic interpretability share a common objective: replacing a pre-trained model (the teacher) with a more compact algorithm (the student) that preserves relevant aspects of its behavior. Distillation typically targets input-output agreement, while mechanistic interpretability instead focuses on creating compact representations that are causally faithful to a model's internal computation, often on a restricted set of inputs. Prior work suggests that constraining a student to reproduce aspects of a teacher's internal computation can improve distillation. We pursue this direction and build on recent advances in causal abstraction that provide differentiable objectives for measuring whether a smaller model constitutes a causally faithful abstraction of a larger one. We use these objectives to develop Causal Abstraction Distillation (CAD) and its practical variant CAB, which trains student Transformers to match not only the teacher's input-output behavior, but also its responses to interventions and perturbations applied to inputs and intermediate activations. This interventional training encourages the student to preserve the teacher's computational structure rather than merely imitate its predictions. Across encoder and decoder-only models, we find that CAB improves intervention-response robustness in several settings while maintaining competitive task performance. CAB can also improve generalization to related datasets, with benefits extending to 14B-to-3B Qwen2.5 distillation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.