acceptodds
Under review as a conference paper at ICLR 2027

Delta-GRPO: Integrating OPD-Style Policy Shift into Reinforcement Learning

Abstract

On-policy distillation (OPD) transfers teacher knowledge through token-level feedback on student-generated responses. Recent work shows that the change between a teacher’s policies before and after reinforcement learning can also provide a useful learning signal. In contrast, group relative policy optimization (GRPO) assigns the same outcome-based advantage to every token in a response, without explicitly using token-level policy changes for credit allocation. We propose a method that combines OPD-style guidance with GRPO’s outcome-based optimization, using the training policy’s accumulated shift to guide token-level credit allocation during reinforcement learning. Across eight mathematics benchmarks, the method outperforms the evaluated RL and OPD baselines with Qwen3-4B and Qwen3-8B. It improves macro accuracy over GRPO by 13.59 and 12.04 percentage points, respectively, and over SC-GRPO by 8.90 and 7.89 points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.