acceptodds
Under review as a conference paper at ICLR 2027

JUDGEGYM: VERIFIER-ANCHORED ENVIRONMENTS FOR ATTRIBUTING LLM JUDGE FAILURES

Abstract

A key strength of LLM judges is that they scale evaluation. A judge scores answers no human has time to check, and the same judge serves as the reward model a policy is trained against. However, we find that the labels used to train and test judges are a source of blind spots. Judges are graded on preference labels, and a preference label says which answer was better and nothing about why, so a judge that misses a flipped sign scores the same as one that favors the longer answer. We identify two failure modes of this practice. First, an asserted label cannot attribute a failure to a mechanism, so training judges harder raises aggregate scores without removing the flaw. Second, the oracle that grades answers is lenient by design, and its tolerance silently deletes the test cases that would expose scale errors. We introduce JUDGEGYM, which builds judge tests from benchmarks whose expert written (gold) answers come with the short program that computes them. JUDGEGYM corrupts a gold program in one known way, re-executes it, and keeps the pair when the benchmark’s matcher says the answer changed. Both right verdicts are therefore computed rather than asserted. Each kind of corruption tests one invariant, and a control that changes nothing must always be discarded. Five kinds of change over 7,095 FinQA programs give 24,875 verified pairs, about a quarter of them with the sign flipped, and the control is discarded all 7,095 times. The benchmark’s matcher treats 0.125 and 12.5 as the same percentage. That leniency hides 60% of the changes to a constant’s scale, and a strict matcher recovers 99%. We prove that this hidden share caps what any model trained against the matcher can gain, at any sampling budget. Since every right response is known, the pairs are an exact training reward as well as a test. We compare JUDGEGYM with nine judge and reward-model benchmarks, specify a one-GPU training recipe and a finance-to-tax-law transfer experiment, and release the engine, the pairs, the harness, and every number as a recomputable artifact. An executed label, not a larger judge, is therefore what makes a judge’s failure attributable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.