acceptodds
Under review as a conference paper at ICLR 2027

When Students Fail to Learn from Teachers: Learning Dynamics of On-Policy Distillation

Abstract

On-policy distillation (OPD) has become a core technique for post-training large language models, enabling students to learn from dense token-level supervision on their own generations. However, even with a capable teacher, OPD training can be unstable, leading to excessive response length, repetitive generation, and severe performance degradation. One example is spurious agreement, where the student maintains high overlap with the teacher at the states visited but produces repetitive generations. Although such failures have been observed, the mechanisms behind these failures remain poorly understood. To address this gap, we investigate OPD through the lens of learning dynamics, tracing three mechanisms from local analysis to subsequent on-policy sampling. We first identify states where sampled-token OPD updates fail to correct teacher–student mismatch and characterize them using the student's mass on the teacher's top- set. We then show that learning from the teacher is difficult at mismatched states and that their gradients can also interfere with learning at other tokens. Finally, we study how continued training on mismatched states influences on-policy sampling, suppressing the student's current modes and leading the student toward states where the teacher is unreliable. These findings help explain the benefits of supervised fine-tuning (SFT) initialization and motivate targeted correction of states where student and teacher are mismatched. We introduce state-dependent KL routing, which applies forward KL when student coverage of teacher-preferred tokens is low or a substantial fraction of teacher probability mass is underestimated. Across three Qwen3 model scales and both Base and SFT initialization, our method improves mean Avg@8 by 1.38–2.50 percentage points and mean Best@8 by 1.01–6.24 points over sampled-token OPD on mathematical reasoning benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.