Distill Where It Matters: Selective On-Policy Self-Distillation
Abstract
On-policy self-distillation (OPSD) provides dense supervision without a stronger external teacher by conditioning a frozen teacher initialized from the same model on the reference solution. However, standard OPSD distills every example equally, overlooking that privileged information may benefit the teacher differently across problems. We introduce Information-Gain Weighted OPSD (IG-OPSD), which measures the teacher's advantage over the student through their reference-likelihood gap and assigns greater weight to examples where this advantage is larger. Experiments with Qwen3 models show that IG-OPSD consistently outperforms uniform OPSD across four mathematical reasoning benchmarks at both 8B and 14B scales. On the two most difficult benchmarks under uniform OPSD, the 14B model improves Avg@12 by 4.2 points on AIME25 and 5.6 points on HMMT25. Ablation and diagnostic analyses show that the gains depend on matching weights to the correct examples: IG primarily captures the advantage provided by privileged context and directs learning toward harder problems. These results establish sample selection as a key design axis for on-policy self-distillation: the student should learn more from the teacher precisely when privileged information makes the teacher more informative.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.