Efficient RL Selects the Judge's Mistakes: Prompt Selection Under LLM-Judged Rewards
Abstract
Group-based reinforcement learning with verifiable rewards increasingly relies on large language models as judges and data selection methods which aim to spend rollouts on mixed prompts (prompts where the group contains at least one accepted and one rejected rollout). We identify a fundamental incompatibility between these two practices: current selection methods do not distinguish between prompts the policy is genuinely learning from and prompts the judge grades incorrectly. We discover that the filtered set of prompts which the policy has not learned tends to concentrate on those which are plagued by the judge’s mistakes, leading to repeated optimization toward answers the judge wrongly accepts or rejects. We show this across selection methods, model and group sizes, learning rates, and the math and science domains. With an exact verifier, these methods deliver their benefits; however, when a judge is used as the verifier, most of their advantage vanishes. In fact, up to a quarter of rollout groups train the policy toward the wrong behavior (versus 5% under uniform sampling), and the number of prompts the policy answers wrongly but the judge accepts increases two to four fold. As these prompts are revisited, each repeated rollout contributes less learning and more exploitation. In mathematics, held-out accuracy shows no consistent change since learned errors are largely problem specific; on the other hand, the policy uses more than double the compute it needs with an exact verifier to achieve the same mastery. On rubric-based science, judge errors appear more stylistic and exploitation transfers to out-of-distribution evaluation, costing 10 to 11 points on MMLU-Pro STEM that are not recovered within training. This loss, too, is selection amplifying the judge’s (stylistic) errors, through the composition of the selected groups. We propose two simple changes to mitigate this issue: bound how much of the rollout budget any prompt may absorb, and use independent verification only on prompts that exceed this limit. Together, they use about 4% additional verifier calls and remove most of the excess exploitation. We also propose importance-weighted selection, which draws prompts in proportion to how likely the selector is to keep them and reweights each by the inverse of that probability, so the expected update equals uniform sampling’s; applied to GRESO, it recovers 9 of 10 lost MMLU-Pro STEM points. Our findings call for reevaluating claims of efficiency for RL data selection under verifiers used in practice and for designing prompt selection and verification in tandem.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.