Faithful Pruning: Preserving Predictions and Interchange Intervention Responses
Abstract
Post-training pruning usually makes a network smaller while preserving its outputs. A smaller network would also make mechanistic interpretability studies cheaper, but these studies depend on more than the outputs. They apply interchange interventions to internal states and measure how the outputs change, and a pruned network that matches the original's predictions can still respond differently to these interventions. We study feed-forward (FFN) channel pruning that aims to preserve both the predictions and these responses. Removed channels are replaced by constants that fold exactly into the layer's bias, so the pruned layer is physically smaller. We introduce a paired-response objective that penalizes the pruned network's output error without the intervention, its output error with the intervention, and its error in the change the intervention causes. A greedy selector fits this objective from a small calibration set and can prune a single layer or every FFN of a network. Against an otherwise identical objective without the paired term, pairing lowers intervention-response error by 7.8% to 19.6% in all six single-layer GPT-2 language settings we test, carries over unchanged to Pythia-410m, and persists when the retained weights are also trained. When every FFN of GPT-2 is pruned, pairing helps on pronoun continuation and gives mixed results on indirect object identification. Without the intervention, the effect of pairing on output error varies across settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.