Posterior-Weighted On-Policy Distillation: Learning When to Trust the Teacher
Abstract
On-policy distillation (OPD) trains a smaller student on its own rollouts, using token-level feedback from a stronger teacher. Yet it trusts the teacher equally on every rollout, and the task reward, when used, enters only as a hand-tuned additive term that competes with distillation. We ask when teacher guidance should be trusted, and answer with a derivation rather than a heuristic: under a Boltzmann model of trajectory optimality, the principled per-trajectory weight is the group-normalized posterior over group-relative advantages. Posterior-Weighted OPD (PW-OPD) thereby demotes the reward from a competing objective to a modulation of a fixed distillation budget—one new temperature, used untuned—and recovers standard OPD exactly when the reward carries no within-group information. We verify the premise directly: student–teacher divergence on high-advantage trajectories is that on low-advantage ones, in both the short- and long-context regimes. Distilling a 27B teacher into a 9B student, we find that in the long-context thinking regime the additive objective drives AIME accuracy from the untrained student's 0.90 down to 0.75–0.82; a reward-free control attributes the damage to the additive term. Posterior weighting preserves and marginally improves the student (0.90–0.92), and the same ordering—posterior above additive above uniform—replicates in a routed multi-teacher extension with a different student and five domain benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.