Masked Self-Distillation: Letting Reasoning Models Keep Thinking
Abstract
On-policy self-distillation (OPSD) has emerged as a post-training paradigm for LLMs. However, we show that naive OPSD degrades reasoning models on tasks that demand exploration and reflection, such as competition mathematics. We trace this degradation to a context asymmetry: exploration and reflection are redundant to a solution-aware teacher, so uniform distillation causes the student to lose both abilities. Specifically, on identical student prefixes, the privileged teacher assigns far lower probability than the student to a sparse set of epistemic markers linked to exploration and reflection, such as “Wait” and “perhaps”. We propose masked self-distillation, which withholds the distillation loss at marker positions and retains it everywhere else. By letting reasoning models keep thinking, the masked objective reverses the degradation in every tested configuration and outperforms SFT and GRPO. More broadly, hindsight is a poor guide to exploration: privileged self-distillation should transfer its preferences selectively, not uniformly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.