acceptodds
Under review as a conference paper at ICLR 2027

Does self-distillation improve reasoning?

Abstract

In self-distillation, a teacher is constructed by prompting a student model with privileged information or environmental feedback, and logit-level feedback is derived from the difference between the two models' output distributions. This paradigm has recently gained traction as an alternative to policy-gradient reinforcement learning, promising dense supervision where reward signals are sparse and resting on the assumption that the self-conditioned teacher approximates the optimal KL-regularized policy. This paper critically evaluates whether self-distillation methods indeed extract useful guidance from privileged information in reasoning tasks or whether their gains are driven by simpler mechanisms. We report four findings: (1) Self-Distillation Policy Optimization (SDPO) only improves performance on LiveCodeBench v6 problems seen during training whereas policy-gradient training on the same reward generalizes to held-out problems; (2) SDPO's reported gains in science reasoning and tool use are largely recoverable through supervised fine-tuning on final answers without chain-of-thought; (3) the gains of On-Policy Self-Distillation (OPSD) in math can be attributed to distilling thinking-mode behavior into the non-thinking student, a mechanism that does not require privileged information; and (4) SDPO and OPSD gradients are not aligned with the policy gradient, contradicting the optimal-teacher assumption. We evaluate across Qwen3 models from 1.7B to 14B and Olmo-3-7B models. Together, these results indicate a need for robust methods that leverage privileged information to improve reasoning performance. Our code is available at https://anonymous.4open.science/r/does-sd-improve-reasoning/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.