More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges
Abstract
Self-reward, self-play, and LLM-as-a-judge pipelines for label-free self-improvement train a language model against its own reference-free judgments, assuming that a model's verdict on a shown answer is a usable proxy for correctness. We show the premise fails structurally: conditioned on a candidate, a reference-free judge scores plausibility, not correctness. This verification asymmetry runs opposite to the usual one, because the verifier is itself reference-free, and it leaves false-positive basins of plausible-but-wrong answers a policy learns to exploit. We measure the failure with a hidden-anchor audit: a held-out exact-match check the judge never sees. On GSM8K with Qwen3 policies in a reasoning-suppressed regime, self-play drives the judge's pass rate from 0.72 to 0.94 while true accuracy stays at 0.20. Unlike the judge biases already documented, these errors are optimized rather than incidental, and a better or a broader judge does not remove them: they transfer across the Qwen, Llama, and Gemma families and to larger judges; a strict three-family ensemble still accepts most of them; recompute prompts leave the false-positive rate at 0.850; and training against the ensemble reward widens the basin. What does fix it is a single measurable bit: whether the judge commits its answer independently of the candidate, a bit detectable from the judge's own solve accuracy before any training. Committing first, with the candidate still in view, holds the false-positive rate on wrong answers at 0.044. Used as the training reward with the candidate withheld, the same independent verdict keeps false positives near zero in every setting we run; in our main setting, true accuracy rises under this reward in every seed while the plausibility reward leaves it flat.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.