Benchmark Reliability: A controlled comparison of human preference, rubric scoring, and LLM judges on EnterpriseVAL
Abstract
LLM judges are validated by their agreement with expert ratings, a practice that treats expert reliability as a fixed property of the raters. We show that it depends substantially on what those raters are asked. On EnterpriseVAL, a benchmark of expert-authored finance and accounting tasks, we hold tasks, outputs and rubrics fixed and vary only what the judgement asks for: blinded preference between an LLM output and the expert reference deliverable; a score on each criterion of the task rubric; and a binary critical-failure gate, where a single disqualifying error fails the output. Expert agreement rises as the claim under judgement becomes more specific: Krippendorff’s α = 0.20 on preference, α = 0.55 on criterion grading, Cohen’s κ = 0.90 on the critical-failure gate. Two consequences follow for validation: a judge’s agreement with experts is uninterpretable without the expert–expert baseline on the same instrument, and preference can favour outputs the gate disqualifies. Reliability trades against discriminative power. The critical-failure gate, our most reliable instrument, separates the evaluated LLMs least, because a criterion every LLM passes, or every LLM fails, carries no comparative information. Reliability and discriminative power are separate claims about an instrument, and a benchmark tuned to maximise the first can end up measuring nothing about the LLMs it exists to compare. Where a single wrong value disqualifies a deliverable, categorical correctness belongs in a benchmark’s results as a claim of its own, reported separately from graded and preference measures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.