Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning
Abstract
Instruction following language models are vulnerable to safety degradation during fine-tuning, even under benign data distributions, and are further susceptible to adversarial updates that systematically erode alignment. Existing defenses offer insufficient protection or impose a steep utility cost, limiting their practical adoption. In this paper, we introduce a training framework that uses frozen activation-based safety signals to adaptively allocate regularization across training examples and individual target tokens, enabling models to remain aligned while preserving task performance. Our framework derives a model-specific intervention direction from pre-generation representations, uses the resulting intervention effects to construct an example-level control signal, and combines this signal with token-specific reference preferences to determine token-level regularization strengths before optimization. Crucially, we demonstrate that pre-generation activations contain linearly predictive signals of safety-critical response outcomes, establishing an empirical basis for proactive, rather than reactive, safety enforcement. Experiments across models ranging from 3B to 8B parameters under both benign and adversarial fine-tuning regimes show that our framework preserves alignment throughout training while maintaining competitive downstream utility. Our results suggest that the safety-utility trade-off commonly associated with alignment-preserving fine-tuning can be substantially reduced through adaptive regularization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.