Beyond Pointwise Signals: Prospective Supervision Weighting in On-Policy Distillation
Abstract
On-policy distillation (OPD) for large language models (LLMs) provides dense teacher supervision along student-generated trajectories, yet existing token-weighting methods judge each position by signals local to the current state. This overlooks a defining property of OPD: because the student generates its own trajectories, learning at one state changes which states, and hence which supervision, it encounters later. Through controlled single-state interventions, we show that updates with nearly identical immediate learning gains can have opposite downstream effects. We therefore introduce Prospective Supervision Weighting (PROW), which scores each token by its immediate learning opportunity together with a prospective term estimating how the teacher-directed correction at that token shifts the student toward continuations offering more learning opportunity. PROW computes both terms from standard OPD rollouts and only reallocates a fixed per-token optimization budget, requiring no additional rollouts, reward models, or learned critics. Across reasoning benchmarks and model scales, PROW improves standard OPD and existing token-weighting baselines, while controlled interventions confirm that its prospective signal captures effects beyond local learning value.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.