acceptodds
Under review as a conference paper at ICLR 2027

Progressive Misalignment Can Emerge from Benign Recursive Self-Distillation

Abstract

Training models on their own outputs has attracted widespread attention as a possible route to further capability gains in AI development. Self-distillation has emerged as a prominent method for this, in which a model is trained to match the token-level predictions it would have produced when given access to additional context, such as a cue, feedback, or instruction. However, the safety properties of this training procedure are poorly understood. In this paper, we present a surprising safety failure of self-distillation. We study recursive self-distillation, where each self-distilled model becomes the teacher for the next round of training. When a safe model is recursively self-distilled on its own safe generations while receiving a fixed non-harmful feedback, it progressively becomes broadly misaligned across harmful queries; we call this failure mode *Progressive Misalignment*. We show this failure mode in both the Qwen3 and Phi-4 model families on SORRY-Bench and StrongREJECT. Progressive misalignment also emerges in a more realistic setup, where the fixed cue is replaced by feedback from a simulated user with benign intent. Using a gradient-based data selection approach, we are also able to find benign prompts that lead to progressive misalignment when applying recursive self-distillation, showing that harmful training prompts are not necessary for this failure mode to emerge. These results suggest that safety might degrade unintendedly across rounds of self-improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.