Stronger, Yet Not Better: Demystifying Self-Distillation Initialization
Abstract
On-policy Self-Distillation (OPSD) efficiently trains LLMs by providing dense, token-level teacher supervision on the student's own rollouts. However, this supervision degrades when student errors push the student's policy too far from the teacher's, making OPSD highly sensitive to initialization: the student checkpoint and teacher configuration that keep supervision reliable and learnable. Prior OPSD methods often assume that stronger students or privileged teachers lead to better distillation. We challenge this assumption with hundreds of experiments spanning model scales, SFT checkpoints, and privileged teacher configurations. At the problem level, higher teacher accuracy on student-failed problems predicts larger post-initialization OPSD gains. At the token level, lower teacher–student divergence predicts larger gains when the student is undertrained, while this relationship weakens or reverses with overfitting. These results show that effective initialization depends not only on model strength, but also on teacher complementarity and learnable teacher–student divergence. Guided by these findings, we introduce a divergence-aware SFT warm-up that upweights low-divergence tokens during OPSD initialization, prioritizing easier-to-learn targets. It outperforms standard SFT+OPSD, yielding larger gains over initialization on AIME24 (+11.67 vs. +5.00) and AIME25 (+11.25 vs. +6.25). Overall, our results suggest that future methods should optimize OPSD initialization for both model capability and teacher–student learnability, rather than simply choosing stronger students or teachers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.