From Traces to Judges: Automatic Design of Evaluation Pipeline for LLM-Based Multi-Agent Systems
Abstract
Evaluating multi-agent systems from their execution traces is hard: traces are long, failure modes are diverse, and existing methods rely on fixed, benchmark-specific judges or manual prompt engineering, so each new benchmark needs a new judge. We introduce AutoJudge, a framework that constructs an adaptive pool of LLM judges inside a fixed evaluation graph. Given a trace, an evaluation taxonomy, and an output schema, a MetaAgent writes the roles and instructions of a set of intermediate judges and a final aggregation judge, enabling input-adaptive evaluation without task-specific training or judge design. To handle long traces, we propose a dual strategy combining full-trace evaluation with a summary-based mode and controlled step-level access to the original trace. Across seven benchmarks with heterogeneous taxonomies, AutoJudge matches or exceeds the published benchmark-specific evaluator on 14 of 19 within-benchmark metrics with a fixed, low-cost judging backbone and no benchmark-specific judge design, while specialized judges with stronger backbones remain ahead on several metrics. Results are stable across repeated runs, and the framework runs at inference time without training. Our code is available at https://anonymous.4open.science/r/AutoJudge-B0E0.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.