Why LLM-Judge Agreement Is Overstated, and How to Fix It
Abstract
To show that automatic grading can be trusted, papers report how often two language-model judges gave the same verdict. Between a judge's reply and that number sits code. A parser that cannot read a reply writes in a default. A tie is broken to a fixed side. An answer that refuses, or is too short, is marked wrong before any judge sees it. The same code runs for every judge, so wherever it fires the judges agree by construction rather than by judgement, and the reported agreement is partly its own work. We read 35 public judging pipelines and find such code in 28 of them (80%), a parser default most often; none reports agreement on the rows it left alone. Applying each kind of rule to 38,601 recorded verdicts, from nine judges on six human-labelled benchmarks, we find that length rules raise Cohen's by a median of +0.071 and significantly in 78% of judge pairs, while a parse-failure default moves it either way by up to 0.08 according to which verdict the parser writes. Those same length rules lower agreement with the human labels. Agreement on the untouched rows exposes 87% of the inflated cases and costs nothing to report, so we ask that papers report it, together with the share of rows their code decided.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.