An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Abstract
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher–student settings, with average gains of points in Avg@16. Remarkably, through Pass@ evaluations up to , we found that LSPD better preserves policy diversity by achieving stronger performance as grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.