Directing Safety Signals to the Appropriate Neurons: Multilingual Safety Alignment through Selective Neuron Masking
Abstract
The safety behavior of large language models varies considerably from one language to another, and prior work has traced such responses to a compact group of internal safety neurons. Current approaches to multilingual safety alignment usually formulate their optimization targets behaviorally, through loss terms or reward signals, yet seldom determine which of the model's safety-related internal representations actually absorb this training pressure. To address this gap, we introduce SNPO, a preference optimization scheme that operates at the level of individual neurons for aligning safety across languages. The method begins by locating the safety neurons tied to each language and then divides all languages into two sets according to how much training data each has. Optimization proceeds by alternating between these sets: on any given step, the safety-neuron activations of one set are masked and their weights held fixed, so that the alignment pressure is funneled onto the neurons belonging to the other set. The assignments are subsequently reversed, allowing the neurons of each set to receive focused optimization in turn. Evaluations spanning seven models and eight languages demonstrate that SNPO lowers how often unsafe outputs occur without sacrificing multilingual capability, and the improvements are most pronounced for languages with limited resources. Examining the neurons directly, we find that after SNPO those safety neurons that had responded to only one language begin firing across several. Taken together, our findings indicate that successful multilingual safety alignment depends not merely on carefully crafted behavioral preferences, but equally on directly governing which internal neurons end up receiving the resulting signal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.