acceptodds
Under review as a conference paper at ICLR 2027

Trial by Trace: Benchmarking AI Agent Harnesses for Reliability and Failure Attribution

Abstract

LLM agents are increasingly used to solve real-world problems, and their effectiveness depends on both the underlying models and the harnesses that support their operation. However, most existing benchmarks focus on model capabilities and overall agent performance, with limited attention to the evaluation of agent harnesses. To address this gap, we introduce **TraceTrial-Bench**, which organizes harness reliability into seven components covering context, execution, governance, lifecycle, observability, tools, and verification. We construct 95 expert-level, mechanism-oriented tasks with executable or state-based ground truth to support semantically equivalent comparisons while preserving the native behavior of each harness. We design a **Trajectory Tracing (TT) module** to capture agent execution traces and introduce **a two-layer evaluation framework** that separately assesses agent performance and harness reliability. In addition, we design a **Multi-source Evidence Acquisition (MEA) module** to provide evidence from the agent context, harness, and execution environment for LLM-based evaluation, and develop a **Failure Attribution (FA) module** that uses explicit mechanism chains to support attribution at the run level. Across 4 harnesses, 9 models, and 95 tasks, our evaluation of 3,420 runs yields macro-averaged harness scores ranging from 59.8% to 72.0%, while 31.1% of runs achieve task success despite failing at least one harness criterion. TraceTrial-Bench provides an auditable and reproducible foundation for comparing agent harnesses, diagnosing their failures, and studying how harness design interacts with model behavior. Code is available at: https://anonymous.4open.science/r/Trace_anoy-F901.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.