acceptodds
Under review as a conference paper at ICLR 2027

The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization

Abstract

Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying maps a relationship between the attack success rate and the over-refusal rate, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.