Beyond Verifier Reliability: Information for Policy Updates in GRPO
Abstract
Reinforcement learning with verifiable rewards commonly uses binary correctness feedback, yet its accuracy on current responses need not identify which policy update most improves true correctness. Accuracy aggregates errors, whereas update value depends on their alignment with changes in response probability. To isolate this missing information, we fix the initial policy, verifier, and complete GRPO menu, then construct binary utility worlds with indistinguishable finite-precision reports but different preferred updates at equal KL. The separation persists under specified perturbations and finite-rollout construction. We also characterize the exact linear utility moments sufficient for finite candidate comparisons and give conditions for reusing an initial correctness contrast, with separate estimation and redistribution errors. Two controlled Qwen3-4B experiments examine the consequences for feedback evaluation. The first holds group confusion matrices and centered reward MSE fixed to test whether these summaries determine update outcomes: the resulting policies differ by 8.53 percentage points in average utility. Only a four-position prefix-KL statistic is matched; a complete-history diagnostic finds unequal sequence KL. The second tests whether targeted execution information improves training at a fixed test count: four diagnostic tests outperform four random tests by 3.9 and 7.0 percentage points in two seeds. Together, the results identify information omitted by reliability summaries and conditions under which decision-relevant information can be retained and reused.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.