acceptodds
Under review as a conference paper at ICLR 2027

Dual-Null LoRA: A Training Dynamics Approach to Preserving Safety Alignment

Abstract

The safety alignment of large language models, typically established as part of post-training pipelines, is known to degrade under subsequent fine-tuning. This phenomenon occurs not only when fine-tuning on harmful examples but also when using harmless, domain-specific utility data. We use recent developments in the theory of deep learning training dynamics to understand this phenomenon and derive a mitigation for the special case of LoRA fine-tuning. In a first-order approximation of fine-tuning dynamics, gradient steps on potentially harmful training examples change the probability of obtaining an aligned response to any other prompt. This change is mediated by the empirical neural tangent kernel (eNTK), which couples each training example to each observed prompt. We introduce dual-null LoRA, a standard LoRA adapter built to eliminate this coupling: Our adapter is initialized in approximate null spaces of both model activations and gradients, and every update during training is projected back into these null spaces. Evaluated on two base models and three fine-tuning datasets (SST-2, AG News, GSM8K), dual-null LoRA effectively keeps harmfulness (HS) and attack-success (ASR) at the levels of the pre-fine-tuning baselines without negatively impacting task learning or general capability, even at 20% harmful training examples.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.