Beyond Decision Flips: An Empirical Study of LLM Safety Judgments
Abstract
Adding noise to a language model's hidden states shows at which layers its safety judgments change, but not why. We study these responses by dividing the same noise energy among tokens in two ways, reversing the noise sign, adding noise only at some token positions, and predicting responses on new inputs. How the noise is divided among tokens determines whether the response depends on the gradient of the clean decision score, and a noise vector and its negative can move a judgment in the same direction. Position interventions link this response to an uneven output change, in which the differences between inputs' outputs shrink much more than the average output. On Qwen2.5-7B, noise before the final prompt position scales these differences to about 0.55 of their clean size, compared with 0.87–0.93 when only the final position is perturbed. At one responsive layer, a model of this change predicts held-out decision margins with RMSE 3.31, versus 7.16 for uniform scaling and 3.28 for a fit made directly to the margins. Models and layers differ in whether a large response comes with this shrinkage, and labeled outcomes show whether the changes help or hurt accuracy. The tests link a layer's response to the output changes behind it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.