From-Base RLVR for Verify-Only Judges: A Rule-Based Reward Turns Base Models into Trajectory Auditors
Abstract
Reinforcement learning from verifiable rewards (RLVR) is the standard recipe for improving how language models solve tasks. We ask whether the same rule-based binary signal can instead train an effective verify-only judge directly from a base model. Casting judging as a balanced accept/reject task and rewarding the model only when its parsed verdict matches the ground truth—no reward model, preference pairs, or human rationales—we optimize with GRPO on two verification-hard domains: auditing GAIA agent trajectories and verifying MuSiQue multi-hop factual answers. The payoff on trajectory auditing is large and replicates across three model families (Qwen3-8B, Qwen2.5-7B, Llama-3.1): e.g. verify accuracy improves (McNemar ), transfers to a real out-of-distribution set of multi-agent failures, and increases with task difficulty. On multi-hop factual verification, a compact 8B judge attains a lower miss rate than four frontier zero-shot judges, albeit with a false-alarm trade-off. We further report two honest negatives—tripling the data gives no gain (the recipe is sample-efficient), and the multi-hop factual payoff does not replicate on one base—that sharpen when the payoff holds. From-base RLVR is a simple, reproducible route to trustworthy verify-only judges.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.