acceptodds
Under review as a conference paper at ICLR 2027

Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant breakthroughs in complex LLM reasoning within verifiable domains, such as mathematics and programming. Recent efforts have sought to extend this paradigm to open-ended tasks by employing LLM-as-a-Judge to provide sequence-level rewards for policy optimization. However, such rewards are inherently sparse, failing to provide the fine-grained supervision required for generating complex, long-form trajectories. Furthermore, existing approaches treat the Judge as a black-box oracle, using only its verdict while discarding the rich intermediate feedback signals encoded in it. To address these limitations, we introduce Grad2Reward, a novel framework that extracts dense process rewards directly from the Judge's inference process via a single backward pass. Notably, Grad2Reward adopts a self-judging mechanism, allowing the policy to improve through its own evaluative signals without reliance on superior external Judges. To effectively leverage these dense rewards, we further introduce Token Advantage Policy Optimization (TAPO), which enables fine-grained policy updates. Extensive experiments demonstrate that Grad2Reward achieves superior performance and training efficiency across diverse open-ended tasks, affirming its effectiveness and broad applicability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.