acceptodds
Under review as a conference paper at ICLR 2027

Attribution-guided Pruning for Circuit Discovery and Targeted Correction in Language Models

Abstract

The parameters of Language Models (LM) encode knowledge acquired during training and jointly support many behaviors at inference. Identifying the components associated with a particular behavior provides insight into its implementation, while modifying them offers targeted control but can affect general capabilities. Circuit discovery and model correction address these respective goals. We connect them through attribution-guided pruning, which uses attribution methods to score neurons and individual weights. For circuit discovery, we retain highly attributed components. For correction, we compare component attribution on behavior-eliciting prompts and diverse reference text, then prune components more strongly associated with the target behavior. We evaluate LRP and one-step gradient saliency, with Wanda as a pruning baseline where applicable. On Indirect Object Identification, LRP yields the sparsest structured circuits across the three tested checkpoints and remains competitive for individual-weight pruning. In held-out evaluations across nine LM checkpoints from 125M to 8B parameters, attribution-guided pruning reduces the mean toxicity score by 43.2-77.6% on prompts whose unpruned continuations exhibit the target behavior, defined in this study by high toxicity, insult, and obscenity scores. Across three models, it improves response uniqueness by 39.8-86.6% on repetition-eliciting prompts. All selected corrections retain at least 90% of their unpruned HellaSwag and ARC-Challenge normalized accuracy and keep WikiText2 perplexity within baseline. Both LRP and gradient saliency support targeted correction, with gradients yielding the best selected operating point for several models. The results show that the same parameter-level pruning framework can yield sparse task circuits and suppress selected generation behaviors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.