acceptodds
Under review as a conference paper at ICLR 2027

Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals

Abstract

On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood. Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student. To understand this difference, we analyze OPD as an idealized continuous-time dynamical system in the small-learning-rate limit. Our diagnostics link these plateaus to premature learning-signal collapse, where the gradient becomes weak relative to the remaining loss. We further prove a local recovery guarantee for teachers sufficiently close to the initial student in a shared parameterization under regularity conditions, offering a conditional explanation for the success of self-RL teachers in our experiments. Across runs with and without loss plateaus, we observe small relative parameter changes (0.025–0.098%) and high similarity between the student's representations before and after OPD (linear CKA across layers). These observations suggest that limited representation adaptation may contribute to learning-signal collapse, a hypothesis that remains to be tested.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.