Mixed-Policy Distillation: Bringing Teacher and Student onto Common Ground
Abstract
Distillation involves both the student and teacher policies, but trajectories aligned with one are not necessarily suitable for the other. In language model post-training, teacher-generated trajectories can depart from the student's inference distribution, while on-policy distillation (OPD) can expose the teacher to unfamiliar student-generated states and weaken its guidance. We introduce Mixed-Policy Distillation (MPD), a post-training framework that constructs trajectories for compatibility with both teacher and student. MPD combines a student-generated prefix with a teacher-generated continuation through an uncertainty-guided handoff and distills the teacher's token distributions across the full mixed trajectory. In this design, the teacher helps shape subsequent states while its guidance remains anchored in the student's own reasoning. Empirical analysis shows that, after the handoff, mixed trajectories achieve lower joint perplexity and stronger next-token agreement than trajectories generated entirely by either policy alone. Experiments across mathematics, science, and tool use on various Qwen and Gemma models demonstrate improved performance over both off-policy and on-policy counterparts with both externally trained teachers and privileged-information-conditioned self-teachers. These results identify mixed-policy trajectory construction as a promising design choice for distillation-based post-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.