acceptodds
Under review as a conference paper at ICLR 2027

Faithful Pruning: Preserving Predictions and Interchange Intervention Responses

Abstract

Post-training pruning usually makes a network smaller while preserving its outputs. A smaller network would also make mechanistic interpretability studies cheaper, but these studies depend on more than the outputs. They apply interchange interventions to internal states and measure how the outputs change, and a pruned network that matches the original's predictions can still respond differently to these interventions. We study feed-forward (FFN) channel pruning that aims to preserve both the predictions and these responses. Removed channels are replaced by constants that fold exactly into the layer's bias, so the pruned layer is physically smaller. We introduce a paired-response objective that penalizes the pruned network's output error without the intervention, its output error with the intervention, and its error in the change the intervention causes. A greedy selector fits this objective from a small calibration set and can prune a single layer or every FFN of a network. Against an otherwise identical objective without the paired term, pairing lowers intervention-response error by 7.8% to 19.6% in all six single-layer GPT-2 language settings we test, carries over unchanged to Pythia-410m, and persists when the retained weights are also trained. When every FFN of GPT-2 is pruned, pairing helps on pronoun continuation and gives mixed results on indirect object identification. Without the intervention, the effect of pairing on output error varies across settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.