acceptodds
Under review as a conference paper at ICLR 2027

Divergence-Routed On-Policy Distillation for Language Models

Abstract

On-policy distillation (OPD) transfers knowledge from a teacher model to a student along student-generated trajectories, typically by minimizing reverse Kullback–Leibler (RKL). However, RKL can provide weak corrective signals for teacher-supported tokens to which the student assigns very low probability. This raises two questions: which positions are most consequential for the behavioral gap between the student and teacher, and what additional supervision is needed at these positions? Through cross-sampling, we find that divergence-based routing more effectively pinpoints positions contributing to the teacher–student performance gap than entropy-based. Further analysis reveals teacher-mode under-coverage at high-JS positions, with standard OPD increasing the large teacher–student probability gaps. Motivated by this findings, we propose Divergence-Routed On-Policy Distillation (DrOPD). DrOPD retains RKL supervision at all positions, and selectively adds forward Kullback–Leibler (FKL) at positions with the high Jensen–Shannon (JS) divergence. Experiments across six math benchmarks show that DrOPD consistently improves over standard OPD and entropy-based routing. DrOPD also reduces the severe under-coverage cases by 37.9% relative to OPD, supporting targeted FKL as an effective complement to OPD.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.