MIRA: Mined Interaction Replay for Agent-Evaluation–Measuring Policy Adherence Beyond the End State
Abstract
Benchmarks for tool-augmented conversational agents check task completion against the final environment state and tool correctness against the action trace. Policy adherence—verifying identity before acting, honoring a required confirmation, not claiming actions never taken—is not checked directly but proxied from the same end-state check. We show this proxy is unsound. Regrading -bench's own released transcripts on behavior, with agents, scenarios and tool traces held fixed, drops all four of its agents from the 100% policy adherence its native grader reports to 69.7%–76.3%: the proxy and the behavior disagree within the same run. We introduce MIRA (Mined Interaction Replay for Agent-Evaluation), which measures policy adherence directly. MIRA mines workflow graphs and ordering policies from real customer-service dialogues, synthesizes multi-turn scenarios from them, and scores each full transcript—tool calls, ordering and policy compliance along the path taken—with an LLM judge validated against double-annotated human labels, alongside a provider-disjoint second judge. Across conversations with five frontier models, policy adherence averages just 28.9% and spans 12.2%–40.9% across models, and 27.1% of runs that pass both the task-completion and tool-correctness checks still violate policy. The model ordering is identical under all three frontier judges. Reliability is also worse than single-trial scoring suggests: no model passes the same scenario in all three trials more than 2.7% of the time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.