AgentEx-Bench: A Step-Level Benchmark for Evaluating Explanation Quality in Agentic Trajectories
Abstract
Large language model agents increasingly expose intermediate plans, evidence, calculations, and tool decisions, creating an opportunity to evaluate how their observable behavior supports task outcomes. However, existing agent evaluation remains centered on task success, action correctness, or domain-specific process checks. These metrics do not determine whether an emitted trajectory provides a well-grounded, specific, and diagnostically valuable explanation of agent behavior. We introduce AGENTEX-BENCH, a benchmark and measurement framework for evaluating observable agent trajectories as auditable artifacts. AGENTEX-BENCH standardizes heterogeneous reasoning and tool-use trajectories and evaluates key steps along five complementary dimensions: validity, task-grounded faithfulness, specificity, diagnostic utility, and evidence alignment. The framework combines source-aware outcome verification, provenance-aware annotation, controlled perturbation tests, and uncertainty-aware comparison across four task families and 15 systems. Our analyses show that task success and trajectory explanation quality provide related but non-equivalent views of agent behavior. Dimension-level profiles reveal distinctions that aggregate scores can obscure, while matched human-reference analyses show substantial uncertainty in fine-grained model ordering. These findings support treating observable trajectory quality as a distinct evaluation target that complements task success and process correctness for agent auditing, diagnosis, and oversight.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.