acceptodds
Under review as a conference paper at ICLR 2027

OPD: Overlap-Gated On-Policy Distillation on Reliable Prefix

Abstract

On-policy distillation (OPD) supplies dense token-level supervision by aligning a student to the teacher's next-token distribution on its own generated rollouts. Existing works mostly implement top- OPD, applying reverse Kullback-Leibler () divergence between the teacher's and the student's renormalized distributions over teacher top- support. However, the teacher scores student-generated rollouts that are inherently off-policy for itself, leaving the reliability of its next-token guidance unclear. Restricting the divergence to the teacher top- support also discards the remaining mass outside that support. These gaps leave two questions implicit: 1) which tokens to align, and 2) how to align those tokens. To answer these questions, we propose Overlap-gated On-Policy Distillation (OPD). As for 1), we quantitatively analyze the reliability of teacher supervision and find that alignment on the student's trajectory prefix is more reliable than alignment on the full context. Therefore, our OPD is applied only to the 20% prefix tokens of each student rollout. As for 2), we theoretically characterize two problems of top- OPD: the Teacher-support Uncovering and the Dominant-Mode Mismatch. Based on these analyses, we propose two support strategies: the Reset Support}strategy attaches a reset atom to every token, pooling mass outside the teacher top-, and the Support Convergence strategy converges the support to the already aligned intersection. Our OPD employs the designed overlap-gated -divergence to interpolate between mode-covering and mode-seeking for per-token alignment, where the overlap is used as a gate to switch the support and adaptively modulate . Extensive experiments on seven common benchmarks demonstrate the superiority of OPD, improving average performance by up to 5.29 points compared to OPD.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.