Beyond the Verdict: What Accuracy Misses When Choosing LLM Judges for Agent Evaluation
Abstract
A coding agent can skip a required action and still claim that it completed it. A judge may accept that claim even though the trajectory shows otherwise. More broadly, when one judge looks better than another, a single accuracy score does not tell us why. The difference may come from the decision cutoff, from how the judge handles misleading trajectory content, or from useful information that never reaches its final verdict. We introduce the *Verdict Audit* to separate these possibilities. We apply it to two practical judge-selection settings. For Qwen3-8B and Qwen3-14B, a large accuracy gap mostly disappears after adjusting the PASS/FAIL cutoff, but the smaller model remains more vulnerable to false completion claims when reasoning is enabled. For Qwen3.6-27B and Gemma 4 31B, Gemma 4 makes better output-level rankings, while Qwen3.6 contains useful error information internally that its final verdict does not use. Reading that information narrows the gap, but the fixes we test do not reliably correct the errors. The result is simple: choosing a judge requires understanding why its decisions differ, not just which model has the higher score.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.