When Mixing Beats Either Policy: Rethinking Rollout Sampling in Large Language Model Distillation
Abstract
On-policy distillation (OPD) trains a student language model on its own rollouts while using a teacher's logits as supervision. However, sampling entirely from the student has two shortcomings: the teacher's guidance becomes unreliable on student-generated trajectories that deviate from the teacher's distribution, and on hard problems the student rarely produces correct paths for the teacher to reinforce. Sampling from the teacher instead incurs high variance in the gradient correction and risks catastrophic forgetting. In this paper, we prove that neither is optimal: the gradient variance of OPD is strictly convex in the student–teacher mixing strength, and its minimum lies strictly between pure student and pure teacher sampling. Building on this theoretical result, we propose MixOPD, which replaces the on-policy sampling distribution with a geometric student–teacher mixture. This method includes two variants: a fixed mixing strength version and MixOPD-Ada, which adapts the mixing strength at each token using the entropy gap between the two models. We implement both variants natively inside the vLLM inference engine with only 39% wall-clock overhead over vanilla OPD. Experiments on agentic, mathematical reasoning, and code generation tasks across two student sizes show that both variants consistently outperform standard OPD by an average of 4 points, indicating that the choice of rollout sampling distribution is a substantial and underexplored lever for improving distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.