Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
Abstract
On-policy distillation offers dense, per-token supervision for training reasoning models, but it remains unclear when this signal helps and when it hurts. Which teacher should be used, and in self-distillation, which context should serve as the supervisory signal? Answering these questions typically requires costly training runs whose aggregate metrics obscure what happens at individual tokens. We introduce a training-free diagnostic framework that operates per token, per question, and per teacher. We derive an ideal per-node gradient, the update that maximally increases the student's probability of success, and develop a scalable targeted-rollout algorithm to estimate it even for long reasoning chains. The gradient alignment score, the cosine similarity between this ideal gradient and a distillation gradient, quantifies how closely a configuration approximates the ideal signal. In distillation training runs with external teachers, the teacher ranking under the better-performing training configuration largely agrees with the score. Across self-distillation settings and external teachers, distillation guidance is more aligned with the ideal on incorrect rollouts than on correct ones, and the best distillation context depends jointly on the student's capacity and the task, motivating per-task, per-token diagnostics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.