Search-Aware Selection of Verification Protocols for Scalable Oversight
Abstract
Scalable oversight asks weaker judges to evaluate stronger agents, but how reliable a judge is also depends on how the agent spends test-time compute. We study two forms of agent scaling with frozen models: best-of-N search, in which the agent ranks its own incorrect candidates by predicted acceptability without querying or identifying the deployed judge, and longer thinking, in which it reasons with a larger token budget. Across logical and mathematical reasoning and science, we find that judge failures are persistent and problem-specific, so the agent can find them without judge access. Increasing search from one to 64 candidates raises false acceptance by 1.8–5.7x. Longer adversarial thinking also raises it, though on most tasks by less than search at matched token budgets. No single verification protocol stays reliable across tasks, and the cheapest adequate protocol can change the form of verification rather than just its effort. We prove a sharp bound on how much false acceptance can grow as search increases. Building on it, we introduce SARCO (Search-Aware Risk-Controlled Oversight), which calibrates at a small search budget and uses statistical tests on held-out calibration problems to return the cheapest protocol that keeps both false and honest acceptance within target at a larger deployment budget, or abstains. On disjoint problems at search budgets of 16 to 64, every protocol SARCO returns meets both targets, whereas empirical risk-forecasting selectors fail one or both targets in 8–16% of their choices. This assurance comes at the cost of coverage: SARCO abstains when the tested protocols cannot be certified from limited calibration data, including on science tasks where no tested protocol meets both targets. For longer thinking, protocol choices certified at a low reasoning budget can fail at higher budgets, so they must be recalibrated near the deployment budget. Our results suggest that oversight sufficiency depends on how submissions are generated and selected, not only on the agent–judge pair.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.