acceptodds
Under review as a conference paper at ICLR 2027

When Metrics Mislead: Auditing the Measurement Layer of LLM Agent Evaluation

Abstract

LLM agents are often evaluated by reducing execution traces to labels such as final-answer correctness, tool verification, and fault propagation. These labels are useful, but they are measurements produced by parsers, behavioral detectors, and harness rules. We study what happens when those measurements do not fully capture the trace's behavior. We use GSM8K, a benchmark of grade-school arithmetic word problems, in a tool-use setting where agents can call a calculator whose outputs are occasionally corrupted. This gives us a clear ground truth for the expected tool result and allows us to compare what the calculator returned, what the model later used, and what the automated detectors reported. We audit three trace-level detectors across six production model families and 30 problems per model. We find that final correctness, visible re-verification, and fault propagation can disagree in systematic ways. An important failure mode is the silent tool-output override. In these cases, the calculator first returns a corrupted value, but the model later uses a different, often correct, value without calling the calculator again. The bad value may therefore stop propagating even though a verification detector that looks for explicit re-execution records no visible re-check. Human review of 90 traces shows perfect agreement with the reviewed comparable cases for final correctness (67/67), fault recovery/impact (19/19), and fault relevance (18/18), but only 67/81 agreement for verification. All 14 verification disagreements fall into four recurring patterns: repeated calculator calls from stuck loops, verbal recognition of an incorrect tool result without recomputation, repeated expressions within a single reasoning step, and quiet correction after an incorrect value is observed. We also find that execution limits affect the reported result. In an ablation study across two models, one model ranges from 21/30 to 27/30 correct as step and token budgets change, while another remains between 27/30 and 29/30. These findings motivate measurement transparency: an automated label should be accompanied by trace evidence supporting it, whether the detector's question applied to that trace, whether the run ended early due to an execution limit, and any human corrections to the label. Our trace viewer makes this information inspectable alongside the reported results.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.