acceptodds
Under review as a conference paper at ICLR 2027

Beyond Final Answers: Evidence-Grounded Evaluation of Agent Trajectories

Abstract

Evaluating an agent trajectory is difficult when correctness cannot be determined from a final string alone: relevant evidence may be distributed across the execution trace, task-time skills, generated files, and executable behavior. We study how such evidence should be exposed to an evaluator and whether a single evaluation pipeline can adapt to different diagnostic targets through a declarative output schema. Our pipeline combines a MetaAgent, a small set of task-specific micro-judges, deterministic checks, and a structured OutcomeJudge, and supports both optional evidence retrieval and direct materialization of evidence into the shared context. On the full 203-case SkillTV split, evidence-delivery policy substantially changes evaluation behavior: materializing task-time skills yields much higher pass recall than output-only context (70.21% vs. 30.43%), while jointly materializing skills and outputs gives 59.9% accuracy among the matched Gemini-2.5-Flash configurations. A vanilla Gemini-2.5-Flash judge given the same evidence reaches only 44.6% accuracy but a higher F1 score (42.0% vs. 30.8%), showing that evaluator architecture changes how the same evidence is converted into a decision. On a 52-run collection derived from agent executions on AFTER, our Gemini-2.5-Flash pipeline raises pass recall to 83.7% (90.1% F1), compared with 30.6% (46.9% F1) for our reproduction of JudgeSkill v4 with Claude Sonnet 4.6; with the same Sonnet 4.6 backbone, our pipeline reaches 53.1% (69.3% F1). Finally, we apply the same core evaluation architecture to Who&When and TRAIL using benchmark-specific output schemas. On the handcrafted Who&When split, it matches TraceElephant in agent attribution accuracy (56.9%) while improving step accuracy from 1.72% to 12.07%, although agent attribution on the algorithm-generated split remains below both baselines; on the TRAIL GAIA split, a larger micro-judge budget raises Category F1 from 15.40% to 27.75%, exceeding the native TRAIL judge's 24.16%, although fine-grained localization remains substantially harder. The results separate three factors that are often conflated in agent evaluation: which evidence is available, how it is exposed and aggregated, and how evaluation criteria are specified across tasks. Code and data are available at https://anonymous.4open.science/r/codegen-agent-eval-6744.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.