acceptodds
Under review as a conference paper at ICLR 2027

-OPSD: Sparse, Stable and Speedy On-Policy Self-Distillation

Abstract

Improving LLM reasoning often requires costly demonstrations, rollouts, or strong external teachers. On-policy self-distillation (OPSD) uses the same model conditioned on privileged context as its teacher, but dense supervision is costly and OPSD sometimes exhibit unstable training. We propose -OPSD, a sparse, stable, and speedy framework combining state selection, stable optimization, and target smoothing. Viewing the teacher as an implicit reward-guided Doob- transform, we define state importance as action-wise advantage variance and derive a low-cost proxy bounded above and below by this variance. We show that SGD enables sparse weight updates, reducing parameter displacement and improving stability over momentum-based optimizers. At selected states, we smooth targets by mixing the teacher distribution with the initial student distribution as a fixed reference policy. Across model families and sizes, mathematical reasoning benchmarks, and multi-task learning, -OPSD improves performance at equal training steps and remains stable over extended training. It also enables memory-efficient training and, at matched effective batch size, reduces training time by up to 13 relative to GRPO and 3 relative to OPSD baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.