Delta-GRPO: Integrating OPD-Style Policy Shift into Reinforcement Learning
Abstract
On-policy distillation (OPD) transfers teacher knowledge through token-level feedback on student-generated responses. Recent work shows that the change between a teacher’s policies before and after reinforcement learning can also provide a useful learning signal. In contrast, group relative policy optimization (GRPO) assigns the same outcome-based advantage to every token in a response, without explicitly using token-level policy changes for credit allocation. We propose a method that combines OPD-style guidance with GRPO’s outcome-based optimization, using the training policy’s accumulated shift to guide token-level credit allocation during reinforcement learning. Across eight mathematics benchmarks, the method outperforms the evaluated RL and OPD baselines with Qwen3-4B and Qwen3-8B. It improves macro accuracy over GRPO by 13.59 and 12.04 percentage points, respectively, and over SC-GRPO by 8.90 and 7.89 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.