acceptodds
Under review as a conference paper at ICLR 2027

Randomized Horizons with Centered History for On-Policy Distillation

Abstract

Teacher-logit on-policy distillation learns from a student's own continuations, making response length a joint budget for generation, teacher supervision, and student differentiation. A fixed short horizon saves work but removes late supervision from the training target. We study an alternative that retains the mean of a specified full-horizon semi-gradient while evaluating response tails intermittently. Centered-history OPD combines a continuation decision shared across the batch with a learned prediction of the omitted tail gradient. A weighted loss produces the student update in one reverse traversal, and a contrast between continuation outcomes updates the prediction without computing a separate tail gradient. We characterize the additional covariance, the role of causal centering, and the distinction between expected retained work and runtime. On MATH, the observed mean correctness is 61.662% versus 60.333% for plain OPD with a 0.6B student, and 75.410% versus 74.623% with a 1.7B student. These comparisons use all three available centered-history runs and both available plain-OPD runs at each scale. The results motivate randomized supervision allocation as a practical design direction for on-policy distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.