The Effect Size Is a Free Parameter of the Judge Panel: Buy the Scale, Not the Sample
Abstract
A panel of LLM judges differs from its gold standard in two ways that behave differently. Writing the panel score of system as , the slope is a pure scale error that multiplies every reported effect size and by itself reorders nothing, while the residual can reorder a leaderboard and, by a classical aliasing argument, is beyond the reach of gold-free aggregation. We measure both across three settings and 99,487 judge verdicts from up to 26 judges in 14 model families. The central finding is that is a free parameter of the panel: on identical items under an identical protocol, it runs from under an open-weight roster to under a frontier one on Chatbot Arena, tracking the panel's mean member accuracy, not the task. Judged effect sizes therefore have no unit until the roster is reported. The fix is cheap: two gold-anchored systems with no residual of their own identify , and a closed-form rule says when rescaling helps; two anchors decide its branch, eight buy its full form. Fed an estimated , the rule's mean regret was never positive, and on a miscalibrated panel 32 anchor labels (64 with the debiased four-anchor estimator) beat 512 spent uniformly under prediction-powered inference on effect-size error (CIs exclude a tie). Anchors cannot buy the ranking: the rank-relevant residual clears a conservative matched null on a deepened open-weight-leaning panel (flip risk ) but not on a commercial one at the same depth, where it is of a median gold gap; and it is neither predictable from surface features nor transferable across panels. Code and every judge verdict are released in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.