TraceJudgeBench: A Unified Benchmark for Validating LLM-as-a-Judge on Agent Execution Traces
Abstract
LLM judges are now a standard way to evaluate LLM-based agents, but the methods and the benchmarks used to validate them do not agree with each other. Each defines its own error taxonomy, annotation scheme, and error distribution, so neither a judge's score nor a benchmark's verdict carries over to another. We present TraceJudgeBench, a meta-analysis that harmonizes ten judge-validation datasets and three unannotated trace corpora into one frame: 846 agent execution traces kept in their native formats, with every label mapped into a shared two-level taxonomy of ten fine-grained and five coarse categories. The frame lets existing and new evaluation methods, each with its own taxonomy, be cross-validated against one another and against the benchmarks they come from. We annotate 231 further traces to fill the gaps public sources leave open: API and system failures, environment failures, and clean executions with no error at all. Applying the frame to a prompted LLM, TRAIL, AEGIS, and a decomposed ensemble judge on one backbone shows what the separate benchmarks hide. Three of the four judges are blind by construction to categories their native schemes cannot name. Every judge reports errors on clean traces: false-error rates run from 27% to 74% for the three established judges and reach 100% for the ensemble run without abstention, and the prompted baseline stays in that range across three backbones. False-error also grows with trace length: on clean labels that two human annotators confirmed trace by trace, the prompted baseline and TRAIL rise 44 and 50 points from short to long clean traces, an association that the data cannot yet separate from task type. A judge validated on short traces carries no guarantee on long ones, and a corpus concentrated at one length cannot show it. The traces, taxonomy mappings, judge-output mapping code, and analysis scripts are included as supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.