acceptodds
Under review as a conference paper at ICLR 2027

Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges

Abstract

Multi-judge evaluation increasingly needs calibrated probabilities, not only verdicts. With a small labeled set, a natural practice is to curate: rank judges by individual accuracy and keep the best few. We argue that the labels are better spent to calibrate the panel: model the judges jointly and choose the panel by held-out proper-scoring risk, which may still select a subset, but never by marginal accuracy. We explain why curation appears to work. A conditionally independent aggregator double counts redundant judges, and in a two-block model we prove that this misspecification alone can make an accuracy-curated panel beat the full panel, while the correctly specified full panel beats both. On four labeled pairwise benchmarks with 38–144 judges, removing only the modeled dependence shifts the comparison against the full panel on every benchmark and split, and two pre-registered tests, a controlled simulation and three held-out panels, agree with this mechanism. Under conditionally independent aggregation the five most accurate judges beat the full panel by 0.09–0.13 NLL on both reward-model benchmarks; dependence-aware aggregation shrinks the gap to at most 0.015. Selecting panels by calibrated risk stays within 0.017 NLL of the best accuracy-ranked rule on every benchmark, without choosing the panel size by hand; the aggregator matters more than the panel rule. JudgeBench remains a counterexample: we recommend never pruning by accuracy alone, not keeping every judge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.