acceptodds
Under review as a conference paper at ICLR 2027

Judge Noise, Not Difficulty: Diagnosing Weak Open-Source Verifiers in RLVR

Abstract

Reinforcement learning from verifiable rewards (RLVR) increasingly replaces the gold verifier with a learned or zero-shot LLM judge, especially in verify-hard domains such as multi-hop question answering where exact-match signals are brittle. A common implicit assumption is that such judges are reliable enough: that their per-example verdicts reflect true answer correctness, so that the group statistics driving RLVR are estimated on signal rather than noise. We show this assumption fails badly for weak open-source judges that are used in practice. On a controlled multi-hop QA verification suite, three weak base judges (Qwen3-8B, Qwen2.5-7B, and Qwen3-4B) exhibit per-prompt verdict instability of 39–66%: given the same question-candidate pair, the same judge flips between correct and wrong across repeated samples. A strong commercial judge is stable (0% flips) on the same prompts, confirming that the instability is a property of weak verifiers rather than the task. This instability is not a parsing artifact: one judge parses a clean verdict 99% of the time yet still flips on 66% of prompts. Because variance-driven rollout allocation cannot distinguish judge noise from true difficulty, it mis-allocates sampling budget toward judge-unstable prompts. We further report a negative control: an RLVR policy trained against a same-source judge yields no robust gain even under that same judge (mean +2.6 points, standard deviation 2.0 points across repeated scoring, spanning zero), because single-run judge scores are too noisy to trust. These results argue that judge-noise diagnostics should be a standard prerequisite before adopting an open-source verifier as an RLVR reward.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.