Rules Move Verdicts Mote Than Paraphrases: Controlled Evaluation of LLM-Judge Panels
Abstract
Agreement does not reveal whether language-model judges follow a stated rule, express an interpretation, or have evidence about an answer. We separate these questions through controlled comparisons. On 160 authored cases and four judges, reversing an explicit rule changes 462/469 complete ambiguous judge-case verdicts; paraphrasing the same rule changes 4/469. The assigned-profile contrast is 90.04 percentage points, with completion bounds [83.59, 97.46]. A prospectively fixed extension reuses all cases with concurrent explicit controls. With neither rule supplied, an absence notice raises undetermined responses by 33.59 points over filler and turns the 41 binary-unanimous filler panels, among 128 ambiguous cases, into mixed responses. This does not make every compatible binary verdict wrong. A 96-case wording follow-up finds positive aggregate undetermined differences for every notice against both filler and convention mention. However, convention mention also eliminates binary unanimity, notice magnitudes vary, and one wording induces undetermined responses on 23/128 determinate controls. Exploratory vote analysis shows that broken unanimity need not remove a binary majority. A separate matched-call comparison retains its observed accuracy tie, not equivalence; older elicitation and checking studies retain their registered nulls, uncertainty and harms. These findings concern fixed cases and judge rosters. Replay code and data are available in the https://anonymous.4open.science/r/llm-judge-replay-iclr2027-DE13/README.md.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.