acceptodds
Under review as a conference paper at ICLR 2027

What a Rubric Judge Measures: Two Defects in a Shipped Scorer Change What It Measures, Not How Noisily

Abstract

Every number a long-horizon memory benchmark reports passes through a rubric judge. The LLM-as-judge literature treats that judge as a noisy instrument — measure the disagreement, rotate the judges, average the offsets away — which presupposes that it measures one latent quantity up to noise and bias. A judge can instead be perfectly reliable and measure the wrong thing. We audit one deployed benchmark's shipped scorer end to end, as code and as a measurement instrument, re-running it under controlled substitutions to measure what it refuses and what it grants. Two defects are readable in the released code: the judge prompt's question placeholder is never filled, leaving the rubric's only irrelevance clause anchored to a literal placeholder, and nine of the ten dimension evaluators truncate the rubric's own half credit through an integer cast. Both change the construct rather than the noise. Filling the placeholder collapses one dimension's score by an order of magnitude and reverses the sign of the correlation between answer length and score; removing the truncation flattens the largest apparent context-length effect in the audit. Handed the record containing the gold answer — the most favourable case any evidence layer can produce — the pipeline still credits only 45 to 63% of items, a shortfall traceable to a frontier model's reading step, not the scorer's leniency, while the same judge separately passes a no-model, one-event reader on nearly a fifth of rubric criteria, many of them satisfied by saying nothing. The credit rate then prices an inference auditors want, from a verbatim-containment count to a bound on resolving power: it is licensed only where a scorer never refuses a correct answer, false for every scorer we measure, so tighter rubrics keep the inference valid and make it unaffordable. One preregistered prediction of ours failed in the field's favour. The rule is to audit a judge as code before auditing it as an annotator: both defects here were readable, and fixable, without running a single system.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.