Attainment Is Not Stationarity: Objective and Optimizer Sources of Behavioral Drift after Safety Alignment
Abstract
As large language models (LLMs) are increasingly deployed in real-world applications, maintaining useful behavior while achieving safety remains important. Yet model behavior can continue to change after a safety metric reaches its target. We show that this post-attainment drift has two separable sources: an objective can retain residual preference pressure after the target safety level is reached, and stored optimizer state can induce behaviorally consequential movement even after the current objective becomes exactly stationary. Preference controls and same-checkpoint optimizer-state interventions isolate these sources. Changing residual preference within the attained safe region changes substantive safe completion; removing the stored first moment eliminates silent-step movement, scaling it produces proportional displacement across three model families, and matched-magnitude directions lead to different safety and refusal trajectories. Thus, safety attainment does not imply objective stationarity, and objective stationarity does not imply parameter stationarity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.