acceptodds
Under review as a conference paper at ICLR 2027

From Traces to Judges: Automatic Design of Evaluation Pipeline for LLM-Based Multi-Agent Systems

Abstract

Evaluating multi-agent systems from their execution traces is hard: traces are long, failure modes are diverse, and existing methods rely on fixed, benchmark-specific judges or manual prompt engineering, so each new benchmark needs a new judge. We introduce AutoJudge, a framework that constructs an adaptive pool of LLM judges inside a fixed evaluation graph. Given a trace, an evaluation taxonomy, and an output schema, a MetaAgent writes the roles and instructions of a set of intermediate judges and a final aggregation judge, enabling input-adaptive evaluation without task-specific training or judge design. To handle long traces, we propose a dual strategy combining full-trace evaluation with a summary-based mode and controlled step-level access to the original trace. Across seven benchmarks with heterogeneous taxonomies, AutoJudge matches or exceeds the published benchmark-specific evaluator on 14 of 19 within-benchmark metrics with a fixed, low-cost judging backbone and no benchmark-specific judge design, while specialized judges with stronger backbones remain ahead on several metrics. Results are stable across repeated runs, and the framework runs at inference time without training. Our code is available at https://anonymous.4open.science/r/AutoJudge-B0E0.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.