acceptodds
Under review as a conference paper at ICLR 2027

Why LLM-Judge Agreement Is Overstated, and How to Fix It

Abstract

To show that automatic grading can be trusted, papers report how often two language-model judges gave the same verdict. Between a judge's reply and that number sits code. A parser that cannot read a reply writes in a default. A tie is broken to a fixed side. An answer that refuses, or is too short, is marked wrong before any judge sees it. The same code runs for every judge, so wherever it fires the judges agree by construction rather than by judgement, and the reported agreement is partly its own work. We read 35 public judging pipelines and find such code in 28 of them (80%), a parser default most often; none reports agreement on the rows it left alone. Applying each kind of rule to 38,601 recorded verdicts, from nine judges on six human-labelled benchmarks, we find that length rules raise Cohen's by a median of +0.071 and significantly in 78% of judge pairs, while a parse-failure default moves it either way by up to 0.08 according to which verdict the parser writes. Those same length rules lower agreement with the human labels. Agreement on the untouched rows exposes 87% of the inflated cases and costs nothing to report, so we ask that papers report it, together with the share of rows their code decided.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.