Revisiting On-Policy Self-Distillation from an RL Perspective
Abstract
Reinforcement learning with verifiable rewards (RLVR) gives a long reasoning trace only a sparse outcome reward. On-Policy Self-Distillation densifies this signal by conditioning the model on one of its own successful rollouts and distilling the resulting self-teacher into the unprompted policy, but its gains on thinking models are fragile. In this paper, we show that, even with an ideal teacher, SDPO performs a one-step behavior-regularized improvement on the value of the behavior policy, rather than maximizing the reward of the policy being improved. To correct this, we derive the exact optimum of behavior-regularized reward maximization and show that, for binary reward, it is a tilted mixture, i.e., an arithmetic mixture of the behavior policy and its success posterior. Based on this result, we propose Tilted-SDPO (T-SDPO), which distills this mixture using SDPO's prompted teacher in place of the posterior and reduces to SDPO when the mixture weight is one. Across math and science reasoning with three thinking models, T-SDPO outperforms SDPO and GRPO at an equal number of optimizer updates, improving math avg@12 over SDPO by up to points. It also keeps longer responses and higher pass@12 than SDPO.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.