Same Verdict, Different Check Values: What Field Order Changes in a Structured LLM Judge
Abstract
Structured LLM judges return a verdict together with per-criterion checks, and auditors rely on those checks to diagnose errors. We show that these checks can depend on the order in which the fields are declared in the output schema, a choice that leaves the content of the schema unchanged. The verdict can remain unchanged, and correct, while the checks differ. Holding the instances, rubrics and all other content fields fixed, we move the rating and confidence fields before or after the checks and rerun the judge. Ten pre-specified randomised designs cover eight (model, task) configurations. On claim verification, 28% of instances retain their verdict while changing their checks when the two runs use different field orders, against 5% when both runs use the same order: an excess of 23 pp. Two further judge models give +15.8 and +7.8 pp on the same task. The effect persists under every remedy and decoding change we tested. Reordering other keys of the schema changes the checks as well. Aggregate accuracy conceals the change: on MCJudgeBench, where every constraint carries a human label, mean constraint accuracy differs by at most 0.34 pp between the two orders. Beneath that, which constraints the judge scores correctly changes, and so does the set of constraints flagged for repair. Used as an evaluator, the same judge ranks a different model first on every cross-order pair of runs, against one pair in five within a single field order. These results show that verdict-level agreement does not certify a structured judge's audit record: error analyses, repair lists and rankings built on the judge can change with the field order alone. We recommend that the field order be reported, and that check-level evidence be read against a replicate in the same field order.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.