acceptodds
Under review as a conference paper at ICLR 2027

When to Use the Solution: Preserving Exploration in On-Policy Self-Distillation for Reasoning Models

Abstract

Does knowing the solution always lead to better supervision for reasoning models? On-policy self-distillation (OPSD) uses a solution-conditioned copy of the model to provide dense token-level supervision on student-generated trajectories. However, preferences formed with solution access may not be appropriate at every reasoning state, particularly when the model is still exploring competing continuations. Our experiments show that uniformly applying such supervision can degrade performance in thinking mode. Our token-level analysis further reveals that solution conditioning frequently changes the unprivileged self-teacher's leading candidate at high-entropy , whereas the leading candidate is usually preserved at low-entropy lock \~ states. These observations motivate adapting the supervision source to the uncertainty of the current reasoning state. We therefore propose Fork-Sensitive On-Policy Self-Distillation (FS-OPSD), which separates the decision of from . Specifically, the method consists of two components: (1) is proposed to mitigate exploration inhibition by selectively applying unprivileged self-teacher at forks and privileged self-teacher at locks. (2) is proposed to balance the weights of the selected source, ensuring supervision targets are appropriately distributed. Our evaluation across five thinking models and three mathematical reasoning benchmarks shows gains in averaged avg@16 up to 8.26% over vanilla OPSD and up to 2.36% over corresponding base models. Code is available at https://anonymous.4open.science/r/FS-OPSD-14A3/README.md.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.