Verification as a Policy: Meta-Verification for AI Agents
Abstract
Verifying whether an AI agent has successfully completed a task could require inspecting its execution trajectory, running executable checks, or consulting other models. Which verification action to perform next depends on the evidence collected so far and the resources remaining. We study Meta-Verification, formulating verification as a sequential decision-making problem in which an agent adaptively selects verification actions and decides when to stop under a resource budget. Rather than following a prescribed verification procedure, we propose Meta-Verification, an approach which provides an agent the ability to construct an instance-specific verification procedure and delegate compute for executing it as evidence accumulates. We introduce AgentVerificationBench, comprising 541 best-of-16 coding tasks and 189 best-of-3 scientific tasks, to evaluate verification accuracy, realized cost, and budget adherence. We find that frontier agents combine diverse verification mechanisms, but the benefits of online replanning vary across models, and unmodified agents frequently exceed their nominal budgets. We therefore investigate harness optimization as a means of improving Meta-Verification through experience while keeping model weights fixed. On a held-out science test set, harness optimization reduces mean verification cost from $23.20 to $5.72 and increases budget adherence from 42.3% to 98.7%, while selection accuracy changes from 60.3% to 55.1%. These findings establish sequential verification-policy construction as a measurable agent capability and demonstrate that its resource efficiency can be substantially improved through harness optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.