acceptodds
Under review as a conference paper at ICLR 2027

What Survives Continued Training? A Longitudinal Benchmark for Emergent-Misalignment Defenses

Abstract

Defenses against emergent misalignment aim to prevent narrow risky fine-tuning from inducing broadly misaligned behavior. Yet a defended model may later undergo benign continued training, and a low misalignment score after this training does not by itself establish that the defense remains effective. Benign training can also suppress misalignment without a defense, while changes in response coherence can alter what the evaluation counts. We introduce a longitudinal benchmark that follows defended models and matched undefended references through the same benign continuation across language models, risky training topics, and defense implementations. By tracking their behavior throughout training and examining the generations underlying each score, the benchmark distinguishes a surviving defense advantage from changes in the reference. We find that, within the benchmark, generated-prose continuation substantially reduces measured misalignment in undefended models, leaving limited room to resolve a further advantage. Continuation on mathematical reasoning tasks preserves a clearer comparison, and defended models often retain an advantage; under supervised fine-tuning they do so even when their own misalignment increases, while under reward-model RL no defended increase resolves. For the tested KL-regularized variant on Llama, lowering the coherence threshold reverses its apparent advantage. These results show that defense persistence requires both absolute behavior and matched comparisons, with explicit accounting for response inclusion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.