Self-Distilled Policy Gradient
Abstract
On-policy self-distillation, where a language model conditions on privileged context to supervise its own generations, is a promising source of dense supervision for sparse-reward reinforcement learning. We instantiate OPD as an auxiliary full-vocabulary student-to-teacher reverse Kullback–Leibler divergence (KL) loss. For any fixed sampled prefix and with the privileged branch detached, we show that its student-side gradient is identical to a policy-gradient update whose centered token advantage is a centered log teacher/student ratio. We propose SDPG, a self-distilled policy-gradient framework that combines standard-deviation-normalized group-relative verifier advantages, exact full-vocabulary OPD, and reference-policy KL regularizationing. Positive-advantage gating and a warmup-decay schedule control when the privileged signal is trusted. Empirically, SDPG improves stability and performance over RLVR and self-distillation baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.