acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Verdict: What Accuracy Misses When Choosing LLM Judges for Agent Evaluation

Abstract

A coding agent can skip a required action and still claim that it completed it. A judge may accept that claim even though the trajectory shows otherwise. More broadly, when one judge looks better than another, a single accuracy score does not tell us why. The difference may come from the decision cutoff, from how the judge handles misleading trajectory content, or from useful information that never reaches its final verdict. We introduce the *Verdict Audit* to separate these possibilities. We apply it to two practical judge-selection settings. For Qwen3-8B and Qwen3-14B, a large accuracy gap mostly disappears after adjusting the PASS/FAIL cutoff, but the smaller model remains more vulnerable to false completion claims when reasoning is enabled. For Qwen3.6-27B and Gemma 4 31B, Gemma 4 makes better output-level rankings, while Qwen3.6 contains useful error information internally that its final verdict does not use. Reading that information narrows the gap, but the fixes we test do not reliably correct the errors. The result is simple: choosing a judge requires understanding why its decisions differ, not just which model has the higher score.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.