acceptodds
Under review as a conference paper at ICLR 2027

MI-SD: Mutual-Information-Guided Multi-Reference On-Policy Self-Distillation

Abstract

On-policy self-distillation improves LLM reasoning by aligning supervision with student-generated trajectories. However, single-reference supervision can reduce trajectory diversity and neglect alternative valid solution paths, particularly in complex reasoning tasks. We introduce MI-SD, a multi-reference on-policy self-distillation framework that supervises student rollouts using multiple valid reference solutions. A shared frozen teacher evaluates each student rollout under distinct privileged reference contexts and provides corresponding supervision. What's more, MI-SD can quantify token-level reference sensitivity using the generalized Jensen–Shannon divergence across teacher distributions. At positions where references provide distinct guidance, MI-SD applies MI-gated max-support supervision. We further introduce a canonical-solution anchor to stabilize training. Across Qwen3 models with 1.7B, 4B, and 8B parameters, MI-SD consistently improves average accuracy on AIME24, AIME25, and HMMT25 over OPSD by 1.3, 0.7 and 1.4 percentage points, respectively. These results demonstrate the benefit of structured multi-reference supervision for on-policy self-distillation across the evaluated model scales. Our code will be made publicly available soon.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.