ControlBench: Did the Approval Actually Happen? Process-Grounded Auditing of Human Control over LLM Agents
Abstract
AI agents now issue refunds, send messages, and modify files on a user's behalf, often after the user types a quick "ok". But did the agent wait for that "ok" before acting? A chat transcript cannot answer this: it records what was said, not the order in which things happened, so an agent that acts first and collects approval afterwards looks exactly like one that waited. We introduce **ControlBench**, a benchmark of 1,800 simulated human-agent sessions ( 8,300 executed actions, two simulators) in which every executed action carries a ground-truth label of whether approval preceded it, plus released real-world agent trajectories audited under the same rule. An LLM judge must then answer one question per action: was it approved before it ran? The answer has two parts. The first is the *record*. Judged from transcripts alone, every model we test errs in both directions, the most approval-optimistic judges approving never-approved actions more often than not; given the platform's event stream, the stronger judges' false approvals all but vanish. The second is the *reader*. Even with complete records, judges fail to apply the rule when an approval arrives late, targets a stale version, or has been revoked, falsely approving most revoked-approval executions; such cases never occur in the released sessions, so ordinary accuracy hides them, while a simple deterministic checker over the same records is exact. Gating actions on event-grounded verdicts rather than transcripts would have blocked nearly every unsafe execution in our logs while costing less. Reliable oversight of AI agents therefore needs both: records of what the system did, and a reader that executes the approval rule. We will release the benchmark, its approval semantics, and all per-item predictions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.