Measuring Say-Do Consistency in Tool-Using Agents: A Claim-Execution Traceability Framework
Abstract
Agent evaluation today follows three paths—state-based, text-based, and goal-plan-action alignment—yet none directly compares what the agent claims to have done against what the tools actually recorded. Deployment risk arises precisely from this divergence: user decisions rest on agent statements. We introduce a criterion-construction procedure for measuring this say-do gap: a scenario-agnostic core (traceability) plus a domain-specific layer deriving claim-execution matching rules from the tool registry, business processes, temporal alignment, and exemptions. Instantiated as a paired text-tool instrument, the procedure runs seven models across three insurance write-operation scenarios (1,780 trajectories), with external replication on 1,980 -bench trajectories. Reading at three refinement levels, it confirms 182 over-promises (40.3% of candidates), every one a fabricated capability no tool supports, a pattern state-based evaluation cannot effectively detect. Transferability is supported by an independent team's timed derivation on an unseen domain and by replication on published benchmark trajectories. The contribution is the procedure, not any specific reading. We release the protocol, codebook, and all adjudication data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.