acceptodds
Under review as a conference paper at ICLR 2027

The Right Direction at the Right State: Dissecting On-Policy Self-Distillation

Abstract

On-policy self-distillation (OPSD) offers a route to improved reasoning by letting a context-augmented copy of the model supervise its own generations. What makes this additional context useful for learning? We investigate this question through controlled studies of mathematical reasoning with Qwen3 thinking models. Solutions to other problems and a single fixed reference can produce useful training effects, while target-specific solutions outperform other-problem references at both evaluated scales. The central finding is that useful guidance depends on the student's reasoning so far. Exchanging reference-induced changes in token preferences between attempts at the same problem reduces continuation accuracy across both scales, with the amount of distributional change matched. On shared Base histories, thinking-only use of the trained policy retains about 90% of its observed full-policy gain. On MATH-500, alignment improves correct completion within 4k tokens by 4.30-4.80 points and reduces mean token use under a 32k cap by 10.7-12.5% relative to exchanged guidance. Together, these findings distinguish the reference a teacher reads from the guidance a student can use. Reference content shapes the learning benefit; the student's evolving reasoning determines the usefulness of the induced guidance. This trajectory-conditioned view connects local supervision during training to successful reasoning completion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.