acceptodds
Under review as a conference paper at ICLR 2027

Directing Safety Signals to the Appropriate Neurons: Multilingual Safety Alignment through Selective Neuron Masking

Abstract

The safety behavior of large language models varies considerably from one language to another, and prior work has traced such responses to a compact group of internal safety neurons. Current approaches to multilingual safety alignment usually formulate their optimization targets behaviorally, through loss terms or reward signals, yet seldom determine which of the model's safety-related internal representations actually absorb this training pressure. To address this gap, we introduce SNPO, a preference optimization scheme that operates at the level of individual neurons for aligning safety across languages. The method begins by locating the safety neurons tied to each language and then divides all languages into two sets according to how much training data each has. Optimization proceeds by alternating between these sets: on any given step, the safety-neuron activations of one set are masked and their weights held fixed, so that the alignment pressure is funneled onto the neurons belonging to the other set. The assignments are subsequently reversed, allowing the neurons of each set to receive focused optimization in turn. Evaluations spanning seven models and eight languages demonstrate that SNPO lowers how often unsafe outputs occur without sacrificing multilingual capability, and the improvements are most pronounced for languages with limited resources. Examining the neurons directly, we find that after SNPO those safety neurons that had responded to only one language begin firing across several. Taken together, our findings indicate that successful multilingual safety alignment depends not merely on carefully crafted behavioral preferences, but equally on directly governing which internal neurons end up receiving the resulting signal.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.