acceptodds
Under review as a conference paper at ICLR 2027

THE JUDGE LOTTERY: DIAGNOSING WHETHER COMPARISONS BETWEEN STOCHASTIC LLM JUDGES ARE ADEQUATELY SUPPORTED BY RESAMPLING EVIDENCE

Abstract

LLM judges gate safety decisions and supply RLHF/DPO training signal, yet a single verdict is stochastic – shaped by decoding temperature/seed, presentation order, and which ensemble member is queried. Prior work shows judges are biased; it does not tell a practitioner whether a specific small-sample comparison is adequately supported by evidence or is instead an artifact of a favorable resampling draw. We call this the Judge Lottery and formalize it with a disagreement rate D(K), a matched-pair order-flip rate, and a scoped, single-example planning estimate K min , together with a five-step protocol (calibration, resampling, a degenerate-verdict screen, regime classification, and Resample-Consensus reporting) for diagnosing, not certifying, whether a comparative judge claim rests on adequate resampling evidence. In a direct pilot on two open-weight judges (Qwen2.5-1.5B/3B-Instruct, 30 shared MT-Bench examples, 1,920 calls, valid rate 1.000), matched AB/BA verdicts disagree at 0.744 (1.5B) and 0.662 (3B), a large effect with no detectable temperature variation that persists under deterministic decoding and that further resampling at fixed order does not correct; a cluster-respecting re-analysis finds the two judges’ rates statistically indistinguishable despite each being individually large. Synthetic content-blind controls show low verdict variance can coexist with a stable, non-evaluative heuristic. We support the protocol with a mechanism-level analogue from a concurrently submitted differential-privacy fairness study and map it onto representative LLM-judge failure modes. A verified post-submission extension to a 7B same-family judge and a second benchmark finds order bias shrinking, but not disappearing, with scale, alongside improving agreement with real human preference labels (Cohen’s κ: −0.122 at 1.5B, +0.368 at 7B). All direct findings are scoped to same-family judges, two benchmarks, and 30 examples; the degenerate-verdict screen and RCS’s order-sampling policy remain diagnostics under refinement rather than finished gates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.