acceptodds
Under review as a conference paper at ICLR 2027

On-Policy Distillation Guided by Evolution of Internal Predictions

Abstract

On-policy distillation lets a student generate the trajectories on which it receives dense token-level supervision, adapting the training contexts to its own policy and outperforming off-policy post-training methods on reasoning tasks. This student-centric view motivates us to investigate another source of model-specific information: how student predictions evolve across depth. Inspired by prior work on internal policy changes in large language models and depth-contrast decoding, we propose Internal Guidance Policy Distillation (IGPD), which combines a frozen teacher with the detached change from a calibrated shallow readout to the final policy. At each sampled prefix, the resulting logit offset modifies the teacher anchor to form an adaptive target. Training jointly aligns the shallow readout with the teacher and distills the final policy toward this target. Because the offset is derived from the current model, it evolves throughout training without requiring a separate reference-policy forward pass. Across two strong-to-weak settings and competition-mathematics benchmarks, IGPD consistently improves over matched OPD and achieves the highest macro-average among the evaluated methods. We hope these findings encourage further exploration of internal prediction dynamics in on-policy distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.