Beyond Nonzero Advantages: Auditing Process Rewards in All-Wrong GRPO Groups
Abstract
In Group Relative Policy Optimization (GRPO), binary outcome rewards assign zero within-group task advantages when every sampled response is incorrect. Adding process scores can restore relative preferences, but a nonzero learning signal does not establish that the resulting update will improve task performance. We propose an auditing framework that separates checkable process evidence, policy-conditioned reward fidelity, and actual update consequences. Task-derived contracts provide deterministic checks on structured evidence without treating partial annotations as complete reasoning ground truth. We define group-relative fidelity through alignment between process-induced gradients and an independently estimated outcome-based reference under the current policy. A separate rule-based audit weight scales process advantages after group normalization and permits abstention when the required checks are unavailable. The proposed training mechanism requires neither an additional learned evaluator nor a teacher model; bounded same-policy continuations are reserved for offline auditing. Using supporting-fact selection as a concrete setting, we specify paired-update tests of whether reference fidelity predicts fresh held-out performance when fact-level scoring accuracy is matched. Coverage-only weighting, randomized weights, and update-magnitude controls are included to distinguish informative evidence selection from simply reducing update strength. This formulation makes explicit the gap between restoring a training signal and justifying its use, and provides a falsifiable route to assessing when inexpensive process evidence should influence learning from all-wrong groups.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.