acceptodds
Under review as a conference paper at ICLR 2027

Locating, Removing, and Restoring: Auditing the Modular Account of Safety in Instruction-Tuned LLMs

Abstract

Safety alignment is typically realised through post-training interventions, and a prominent line of work has proposed a modular account in which sparse sets of “safety neurons” mediate refusal. This account carries an operational claim: that such neurons can be located and pruned to remove refusal while leaving broader generative competence largely intact — a claim evaluated almost entirely through attack success rate (ASR). We examine it at three stages: locating the neurons, removing them, and putting them back. Locating the neurons is not a settled procedure: across four open-weight instruction-tuned families (three in the main text), localisation is architecture-dependent, ranging from concentrated mid-layer MLP blocks to structure distributed across attention and MLP subspaces, and the two scoring strategies in common use are not interchangeable — under one of them, one architecture does not de-align at all but collapses into incoherent generation, and a second does so at a higher ratio, which ASR records as perfect safety. Removing the neurons changes more than refusal: measured against random pruning and intrinsic self-variance, next-token distributions, semantic content, and hidden-state trajectories all shift, with harmful-domain contradiction rising 25–29 points over the unpruned baseline; benign-domain effects are weaker and model-dependent, and most deviations are not accompanied by conspicuous surface breakage, so many would not be caught by inspection for obvious degeneration. Putting the neurons back shows that the two effects are not separated under a median split of the criterion's own ranking: reinstating the higher-scoring half of the removed neurons recovers 91–95% of the ASR shift and 84–86% of the contradiction shift together, while the lower-scoring half recovers 8–18% of either, in all three architectures. That the semantic effect reverses at all rules out irreversible damage inflicted by the act of pruning; under this split, neither half recovers refusal without also recovering semantic stability, leaving the modular reading unsupported by the criterion's own ranking. We conclude that pruning-based evidence does not support the modular reading, and that ASR is an insufficient basis for claims about preserved competence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.