acceptodds
Under review as a conference paper at ICLR 2027

RUC-SDPO: Reward-Uplift-Calibrated Self-Distillation Policy Optimization

Abstract

Reinforcement learning with verifiable rewards (RLVR) enables scalable post-training, with Group Relative Policy Optimization (GRPO) widely used to translate verifier outcomes into policy updates. However, GRPO assigns the same group-relative advantage to every token in a trajectory, resulting in coarse credit assignment; its reward-driven update vanishes when all trajectories in a group receive identical rewards. Privileged on-policy self-distillation can provide dense token-level guidance from training-only context, but applying it indiscriminately may impose redundant, noisy, or reward-irrelevant preferences. We introduce Reward-Uplift-Calibrated Self-Distillation Policy Optimization (RUC-SDPO), a reward-anchored extension of GRPO that calibrates privileged self-distillation by the estimated task-level utility of the privileged context. Specifically, RUC-SDPO retains GRPO as the primary reward anchor on every sample, and uses reward-uplift calibration to estimate the verifier-reward uplift of privileged conditioning over deployment-matched conditioning. It maps only positive uplift to a bounded detached weight, and scales token-level distillation accordingly. The method requires no separate teacher model, learned reward model, or step-level annotations. Across mathematical reasoning and tool-call generation with Qwen3-4B, Qwen3-8B, and Qwen3-14B, RUC-SDPO outperforms a strengthened GRPO baseline, attaining the best average accuracy in all nine mathematical reasoning model–benchmark settings and the best score in eight of the nine tool-call settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.