And a Good Judge Too?: Methodological Lessons for Auditing LLM Judges in Rubric-Based Evaluation
Abstract
LLM judges have become core infrastructure in automated AI evaluation, yet their reliability remains underexamined: in our rapid review of 37 recent ACM benchmarks using LLM judge scoring, over 40% report little or no judge analysis. We audit and extend L2-Bench, a rubric-based benchmark for second language learning design. This is a "hard'" domain marked by expert disagreement, where ground truth is contested and LLM judges are most likely to be deployed. Guided by a recent judge-failure taxonomy, we show that L2-Bench's production judge (Claude Sonnet 4.6) is statistically stable, but most importantly we demonstrate that stability alone cannot establish validity because of rater label skew—a problem common to evaluation in “hard” domains. To anchor judge scores to human expertise, we import rater-severity modeling from educational measurement, with a GLMM attributing 10.1% of latent-scale variance to rater severity, which we propagate to provide strong evidence of convergent validity ( vs. human ceiling, overlapping 95% CIs). We argue that valid judge claims in hard domains require both kinds of evidence, and propose the Minimum Viable Judge Assessment (MVJA), a tiered reporting template for standardizing judge analysis across the field.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.