From LLM Judge Scores to Evaluation Design: A Generalizability-Theory Framework
Abstract
LLM judges provide task-specific feedback for selecting responses, comparing models, and training policies. Their scores combine stable differences between responses with variation introduced by the judging procedure. We present a generalizability-based framework that profiles each fixed judge, projects grading precision from a pilot, and checks whether a protocol improves decisions against an external criterion. We instantiate it with twelve open-weight judges grading the same responses on three sources with correctness, rubric-quality, or human-preference references. The crossed design varies two close prompt paraphrases and three generations. We find a small pilot per source predicts generation-noise rankings on separate queries (correlations 0.93–0.95) and approximate grading disagreement. A second grading under unchanged wording improves aggregate ordering point estimates by 0.89–2.95 percentage points. An independent-call diagnostic shows that averaging can strengthen both reference-consistent and reference-opposed preferences. The framework yields judge-specific precision and validity profiles under declared conditions and separates protocol planning from judge selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.