acceptodds
Under review as a conference paper at ICLR 2027

Cover the Teacher, Stay Close to the Student: Mode-Covering Trust Regions for On-Policy Distillation

Abstract

On-policy distillation should preserve teacher-supported alternatives at student-generated prefixes, yet its usual reverse-KL signal is mode-seeking: alternatives assigned little probability by a capacity-limited student receive weak supervision and can disappear. Proximal Target Distillation (PTD) makes mode covering explicit. It projects the teacher onto a student-relative KL ball, preserving teacher-supported modes while bounding target movement, then fits that target with forward KL to train underrepresented tokens directly. On the normalized teacher-top- support, the projection follows a closed-form normalized Lambert- path whose radius bounds the initial conditional-logit gradient. In a misspecified Gaussian study, repeated PTD updates recover substantial probability in a teacher region initially nearly absent from the student. Controls identify forward fitting as the recovery mechanism and show that proximal target geometry controls the recovery trajectory. Across math-only and teacher-routed multi-domain language-model experiments, PTD delivers the highest Macro best-of- point estimate among trained student methods. Its advantage over top- reverse KL widens with the sampling budget and overtakes linear interpolation at the larger budget. The controlled density recovery and the growing multi-sample advantage jointly show how a mode-covering objective preserves useful alternatives that a mode-seeking update can lose.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.