acceptodds
Under review as a conference paper at ICLR 2027

When a Reliability Coefficient Orders References Backwards: Identifiability, a Calibrated Test, and Design Rules for the Human References of LLM Judges

Abstract

LLM judges are validated by checking that they agree with a panel of human raters, and the panel is trusted when its reliability coefficient is high. We show that this can go wrong in two ways. First, a judge can pass validation by fitting one panel. When two independent panels rate the same items, their covariance measures what they agree on: the common signal, which is not necessarily ground truth. Under an additive model, and for a judge whose errors are independent of the reference's, this sets a validity ceiling on how well any judge can agree with a panel through that signal. On SummEval, G-Eval's published correlations with the expert panel (0.455–0.582) exceed the upper confidence bound of this ceiling (at most 0.30). Recomputed from its released outputs, G-Eval correlates 0.50–0.57 with the experts and only 0.02–0.06 with the crowd. Nine recent judges show the same split in 33 of 36 judge–dimension pairs, one of them borderline. Second, the reliability coefficient can rank references backwards. A component shared only within one panel raises its reliability and lowers its fidelity at the same time. In 19 of 28 D3code region pairs, the panel that looks more reliable is the less faithful one. We give a per-panel test for this component, with 4.8% false positives at a nominal 5% and 93.4% power at \tau^2=0.3. Without additional information, no one- or two-way variance-components model of a single panel can separate the component from signal, and the designs that maximise reliability or fidelity cannot be checked. A checkable design costs only O(1/b) in fidelity. Only five of the fourteen released references we survey allow the check. Identifying the common signal from two panels is classical, and normalized cross-replication reliability already reads it as a disattenuated agreement. What is new is the per-panel test and what it reveals about how LLM judges are validated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.