acceptodds
Under review as a conference paper at ICLR 2027

WHICH VERIFIER ERRORS REACH THE GRADIENT? A MOMENT DECOMPOSITION FOR AUDITING GROUP- RELATIVE REINFORCEMENT LEARNING

Abstract

Group-relative RL with verifiable rewards updates the policy with score-weighted contrasts of verifier rewards, yet verifiers are audited with the intraclass correlation (ICC) or effective sample size (ESS), statistics designed for a group mean. We characterize exactly which properties of verifier error reach the update. Casting the verifier-induced perturbation as a U-statistic, Hoeffding’s variance formula decomposes its second moment into three channels: the within-group error variance, which a mean-square-within (MSW) audit estimates; a coupling between squared error and score norm; and a squared reward-hacking gradient, whose bias accumulates across a batch instead of averaging out. The ICC and ESS capture none of them. The decomposition covers length-normalized updates and shared verifier randomness, and yields an unbiased audit statistic. On two Qwen2.5 policies the MSW channel carries 0.90–1.00 of the per-group perturbation, so variance-based audits are a sound default; yet for a format-sensitive checker the systematic channel reaches 0.26 [0.15,0.37] of a 64-prompt batch update (post hoc). In a pre-specified three-seed audit with LLM judges, ESS ranked three verifiers in the reverse order of their gradient perturbation, while MSW ranked them correctly. These results give practitioners a principled, inexpensive audit target.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.