-OPSD: Sparse, Stable and Speedy On-Policy Self-Distillation
Abstract
Improving LLM reasoning often requires costly demonstrations, rollouts, or strong external teachers. On-policy self-distillation (OPSD) uses the same model conditioned on privileged context as its teacher, but dense supervision is costly and OPSD sometimes exhibit unstable training. We propose -OPSD, a sparse, stable, and speedy framework combining state selection, stable optimization, and target smoothing. Viewing the teacher as an implicit reward-guided Doob- transform, we define state importance as action-wise advantage variance and derive a low-cost proxy bounded above and below by this variance. We show that SGD enables sparse weight updates, reducing parameter displacement and improving stability over momentum-based optimizers. At selected states, we smooth targets by mixing the teacher distribution with the initial student distribution as a fixed reference policy. Across model families and sizes, mathematical reasoning benchmarks, and multi-task learning, -OPSD improves performance at equal training steps and remains stable over extended training. It also enables memory-efficient training and, at matched effective batch size, reduces training time by up to 13 relative to GRPO and 3 relative to OPSD baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.