Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
Abstract
On-policy self-distillation (OPSD) improves mathematical reasoning by training a student to match a privileged teacher distribution at each prefix of the student’s own sampled trajectory. In standard OPSD, the teacher is conditioned on a privileged context, typically a reference solution, but the distillation loss is applied along an unverified student rollout. This formulation couples two inputs with distinct roles: the sampled trajectory serves as a supervision scaffold, determining the prefixes at which supervision is applied, while the privileged context determines the teacher distribution at those prefixes. We separate these two factors in a factorial analysis and find that whether the scaffold reaches the correct final answer has a much stronger effect on downstream accuracy than whether the context is correct. When distillation uses unverified scaffolds, the teacher can rely on information unavailable to the student, creating an imitation gap. This gap decreases with larger model sizes, yet OPSD still allocates most supervision to unverified trajectories. In contrast, verified scaffolds provide supervision that the student can realistically imitate, even when the context comes from the student’s own unsuccessful attempt. Based on this finding, we introduce OASIS (On-policy Alignment via Scaffold-Isolated Supervision), which retains the OPSD objective but limits supervision to the student’s own verified trajectories and replaces the written reference solution with a failed attempt generated by the model. This method provides a scaffold within the student’s trajectory distribution, allowing OASIS to use only final-answer labels without requiring worked solutions. Across Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT 2025, OASIS improves over the base model by 3.2 to 3.8 points on average, while the benefit from OPSD drops from 3.05 points at 1.7B to 0.14 at 8B. At 8B, OASIS outperforms OPSD by 3.05 points, showing that restricting self-distillation to verified, on-policy scaffolds maintains its effectiveness as models scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.