On-Policy Distillation Guided by Evolution of Internal Predictions
Abstract
On-policy distillation lets a student generate the trajectories on which it receives dense token-level supervision, adapting the training contexts to its own policy and outperforming off-policy post-training methods on reasoning tasks. This student-centric view motivates us to investigate another source of model-specific information: how student predictions evolve across depth. Inspired by prior work on internal policy changes in large language models and depth-contrast decoding, we propose Internal Guidance Policy Distillation (IGPD), which combines a frozen teacher with the detached change from a calibrated shallow readout to the final policy. At each sampled prefix, the resulting logit offset modifies the teacher anchor to form an adaptive target. Training jointly aligns the shallow readout with the teacher and distills the final policy toward this target. Because the offset is derived from the current model, it evolves throughout training without requiring a separate reference-policy forward pass. Across two strong-to-weak settings and competition-mathematics benchmarks, IGPD consistently improves over matched OPD and achieves the highest macro-average among the evaluated methods. We hope these findings encourage further exploration of internal prediction dynamics in on-policy distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.