acceptodds
Under review as a conference paper at ICLR 2027

Available, Relevant, Authorized? Evaluating LLM Agents with Action-to-Evidence

Abstract

Evaluating LLM agents requires distinguishing what they propose, what policy permits, what executes, and what changes outside the model. Conflating these stages can turn missing evidence into apparent safety or credit recovery after an unauthorized effect. We introduce Action-to-Evidence (A2E), a claim-admission mechanism that determines which conclusions execution records support. A2E binds action, policy, invocation, and effect records to the same execution attempt. Effect claims require consistent witnesses; absence and compliant recovery additionally require complete trajectory coverage and observation closure. Incomplete evidence leaves a consistent claim unresolved; conflicting records are rejected as support for that claim. Across 10,934 controlled conformance queries, A2E matches every preset verdict, while removing individual checks admits 801–2,932 unsupported claims. MCP-COMPOSE, our Model Context Protocol benchmark, spans twelve model endpoints, with policy-linked evaluation on five. Replaying frozen proposals isolates the effects of host decisions on execution. In a separate prescribed-action panel with per-call service resets, three invocation-time controls each record 0/180 denied effects but 174–175/180 compliant recoveries. One denied non-target effect followed by a fallback exposes errors in both target-only and fallback-only scoring. On 120 matched attacked AgentDojo trajectories per control, tool filtering and transcribed Progent rules record no benchmark-defined attack effects but complete 2 and 87 tasks, respectively. Evaluations must report effect prevention, compliant recovery, and task utility separately.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.