A LANGUAGE-MODEL JUDGE’S RESOLUTION IS PREDICTED BEFORE IT IS MEASURED
Abstract
A language-model judge cannot tell every pair of answers apart. Below some difference in quality it is indifferent, and that threshold decides whether it can grade the work in front of it. We register two gates. The first is that the threshold can be predicted before it is measured, and the other is that a second budget can reverse the judge’s preference outright. Judges grade worksheets whose quality is known exactly. A calibration block records the scores a judge actually uses, and from that block alone we predict, and commit with its hash, how often the judge will order fresh pairs correctly at every difference in quality, on three rating scales and six ways of reading its report. Only then is the test block’s seed drawn. Resolution. The prediction held on every graded judge, the largest deviation 0.078 in accuracy against a registered bar of 0.14. The scale a judge is offered is the wrong budget, and no one number drawn from its reports is the right one. The 7B judge asked for a score from 0 to 100 writes nine distinct scores and resolves 5.3 wrong answers in twenty where the scale promises one, and on every scale of every judge the channel it writes fits its accuracy better than the nominal scale does. Reading more of the report lowers the threshold by the predicted amount on five of the six graded judges. Reversal. Give the worse of two worksheets a confident header and a check mark on every line. With no reasoning budget the judge prefers the decorated worse worksheet, only 0.125 of test pairs favouring the better one, and at 1024 tokens of reasoning it favours the better one in 0.95. Measuring what the evidence and the decoration are worth in separate blocks, in which the two never conflict, predicts that reversal, and a second ladder that brackets the crossing locates it, predicted at 72 tokens and observed at 57. On a second vendor’s judge it is predicted at 187 and observed at 166, from that judge’s own calibration. The prediction and the comparison against the nominal scale replicate on a second vendor’s models served through an API, on summaries checked against a passage, which no judge does exactly, and on readers’ reviews scored against the star rating each writer gave. Read to 200 characters, to 400, and whole, the judge’s resolution on those reviews coarsens in that order by the amount each budget’s own calibra- tion predicted. Every bar was fixed before the first of these runs and none was refitted. What predicts a judge’s resolution is its measured channel, of which a scalar budget is one summary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.