acceptodds
Under review as a conference paper at ICLR 2027

Qualifying AI Agents Under Cost-Aware Human Oversight

Abstract

As AI agents move into real workflows, evaluation must answer more than whether an agent can complete a task. Deployment requires determining under what operating conditions a system can meet a specified reliability requirement at acceptable cost, and whether the available evidence supports that operating point. We introduce Reliable Enterprise Agent Deployment (READY), a framework that turns workflow-level evaluation evidence into a statistically supported deployment qualification. READY preserves each workflow’s definition of successful execution while evaluating the complete agent–workflow–oversight configuration. Given a prespecified class of oversight policies, it evaluates their reliability and operating cost, tests them on held-out cases with finite-sample error control, and returns the lowest-cost member of the certified set. The resulting deployment profile characterizes reliability, autonomous coverage, human-oversight burden, and cost rather than reducing a system to a single benchmark score. We instantiate READY in the terminal accept-or-escalate setting on CliniCARE-Bench clinical audit (18 systems, 750 cases) and τ2-bench retail (17 systems, 114 tasks). Across both workflows, autonomous accuracy shows no clear correlation with how well self-reported confidence ranks successes above failures (Pearson r = −0.07 and −0.13). Consequently, systems with similar autonomous performance can require substantially different oversight to meet the same reliability target. On clinical audit, Sonnet 5 and GPT-5.4 differ by only 0.3 points in development accuracy, yet their certified operating points require 12.0% versus 29.3% human review at the same target. Accounting for agent and review costs can further change which certified configuration is preferred. These results show that deployability is not an intrinsic property of an agent or its benchmark score, but of an agent-workflow- policy configuration under explicit assumptions, with routing quality, statistical uncertainty, and cost jointly shaping the reliability-oversight-cost tradeoff.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.