acceptodds
Under review as a conference paper at ICLR 2027

Agreement Is Not Validity: A Preregistered Audit of Human-Anchored LLM-Judge Evaluation

Abstract

Agreement among automatic judges is frequently reported as evidence that an LLM-based evaluator is reliable. Agreement, however, is a reproducibility property of a panel; validity concerns whether its verdicts track an independent target. We preregistered and executed an artifact reanalysis of this distinction using three public, item-level, human-referenced artifacts spanning twelve author-defined artifact–domain cells. It contrasts judge–judge Cohen's κ with the same judges' human-anchored κ and tests transport of a learned mapping from transformed reward scores to agreement with the released reference. The design uses artifact-macro aggregation, equal domain weight, and a single 10,000-draw BCa bootstrap stream. The registered frozen computation returned GRAY. The H2 macro gap was 0.129636 (95% BCa CI [0.109357,0.151594]): distinguishable from zero but below the registered 0.15 materiality margin. Calibration transport showed the same intermediate pattern, with macro ΔECE 0.044192 (95% BCa CI [0.037323,0.070487], Holm-adjusted p=0.0002), below its 0.05 margin; ΔRisk was GRAY-null because only two domains had valid support. Excess agreement relative to the released human reference is measurable and artifact-dependent; it does not identify construct validity. Exploratory exclusion of PPE human ties reduces the H2 macro to 0.073393. Reference noise and label construction limit the interpretation of the gap.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.