Randomized Horizons with Centered History for On-Policy Distillation
Abstract
Teacher-logit on-policy distillation learns from a student's own continuations, making response length a joint budget for generation, teacher supervision, and student differentiation. A fixed short horizon saves work but removes late supervision from the training target. We study an alternative that retains the mean of a specified full-horizon semi-gradient while evaluating response tails intermittently. Centered-history OPD combines a continuation decision shared across the batch with a learned prediction of the omitted tail gradient. A weighted loss produces the student update in one reverse traversal, and a contrast between continuation outcomes updates the prediction without computing a separate tail gradient. We characterize the additional covariance, the role of causal centering, and the distinction between expected retained work and runtime. On MATH, the observed mean correctness is 61.662% versus 60.333% for plain OPD with a 0.6B student, and 75.410% versus 74.623% with a 1.7B student. These comparisons use all three available centered-history runs and both available plain-OPD runs at each scale. The results motivate randomized supervision allocation as a practical design direction for on-policy distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.