Improving the Credit Assignment of Group Relative Policy Optimization via λ-Return Advantage Estimation for LLM Reasoning
Abstract
reinforcement learning from verifiable rewards (RLVR) in large language model (LLM) reasoning. Yet GRPO assigns one outcome-level, group normalized advantage to every token of a response: it cannot tell which tokens actually helped the model reach a correct answer, it pays no price for intermediate verbosity, and its Monte-Carlo estimator carries high variance when responses are long. We introduce λ-GRPO, a practical recipe that repairs GRPO’s credit assignment while keeping its group-relative stabil ity. A lightweight critic network, attached to the policy trunk, is trained on λ-return targets; token-level λ-return advantages are then computed from temporal-difference errors, standardized across the token batch, and com bined with an asymmetric clip-higher surrogate, an explicit per-token length penalty, a KL regularizer to the reference policy, and a λ-annealing schedule that shifts from outcome-level credit early in training to fine-grained token level credit as the critic matures. On GSM8K and MATH with Qwen2.5- 1.5B/7B, λ-GRPO improves accuracy over GRPO, PPO, RLOO, VAPO and DAPO at equal compute, reaching 91.2% GSM8K / 65.8% MATH on the 7B model while shortening the average response from 502 to 327 to kens, and 78.8% / 48.4% on the 1.5B model. Component ablations and bias–variance analysis show that each ingredient contributes consistently, and that value-based token-level credit, not verbosity, is what drives the gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.