acceptodds
Under review as a conference paper at ICLR 2027

From-Base RLVR for Verify-Only Judges: A Rule-Based Reward Turns Base Models into Trajectory Auditors

Abstract

Reinforcement learning from verifiable rewards (RLVR) is the standard recipe for improving how language models solve tasks. We ask whether the same rule-based binary signal can instead train an effective verify-only judge directly from a base model. Casting judging as a balanced accept/reject task and rewarding the model only when its parsed verdict matches the ground truth—no reward model, preference pairs, or human rationales—we optimize with GRPO on two verification-hard domains: auditing GAIA agent trajectories and verifying MuSiQue multi-hop factual answers. The payoff on trajectory auditing is large and replicates across three model families (Qwen3-8B, Qwen2.5-7B, Llama-3.1): e.g. verify accuracy improves (McNemar ), transfers to a real out-of-distribution set of multi-agent failures, and increases with task difficulty. On multi-hop factual verification, a compact 8B judge attains a lower miss rate than four frontier zero-shot judges, albeit with a false-alarm trade-off. We further report two honest negatives—tripling the data gives no gain (the recipe is sample-efficient), and the multi-hop factual payoff does not replicate on one base—that sharpen when the payoff holds. From-base RLVR is a simple, reproducible route to trustworthy verify-only judges.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.