Judge Noise, Not Difficulty: Diagnosing Weak Open-Source Verifiers in RLVR
Abstract
Reinforcement learning from verifiable rewards (RLVR) increasingly replaces the gold verifier with a learned or zero-shot LLM judge, especially in verify-hard domains such as multi-hop question answering where exact-match signals are brittle. A common implicit assumption is that such judges are reliable enough: that their per-example verdicts reflect true answer correctness, so that the group statistics driving RLVR are estimated on signal rather than noise. We show this assumption fails badly for weak open-source judges that are used in practice. On a controlled multi-hop QA verification suite, three weak base judges (Qwen3-8B, Qwen2.5-7B, and Qwen3-4B) exhibit per-prompt verdict instability of 39–66%: given the same question-candidate pair, the same judge flips between correct and wrong across repeated samples. A strong commercial judge is stable (0% flips) on the same prompts, confirming that the instability is a property of weak verifiers rather than the task. This instability is not a parsing artifact: one judge parses a clean verdict 99% of the time yet still flips on 66% of prompts. Because variance-driven rollout allocation cannot distinguish judge noise from true difficulty, it mis-allocates sampling budget toward judge-unstable prompts. We further report a negative control: an RLVR policy trained against a same-source judge yields no robust gain even under that same judge (mean +2.6 points, standard deviation 2.0 points across repeated scoring, spanning zero), because single-run judge scores are too noisy to trust. These results argue that judge-noise diagnostics should be a standard prerequisite before adopting an open-source verifier as an RLVR reward.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.