acceptodds
Under review as a conference paper at ICLR 2027

Review the Window, Not the Step: Oversight of Agents That Choose Their Own Workflow

Abstract

An AI agent that decomposes its own task also decides what a per-step monitor sees. We model this as a zero-sum inspection game: the agent spreads a harmful objective across a workflow of its chosen length, and a principal spends a budget of reviewer reads. If recognition is convex in the share of a plan that a reviewer reads, the agent's best response is to spread harm evenly, and at equilibrium no per-step review policy, however it selects steps, beats uniformly random review; at a fixed audit rate, per-step detection decays as in workflow length. Reading steps jointly at an unpredictable location escapes this decay and dominates per-step review by a factor . We measure the recognition exponent and test the theory on 255 harmful DecomposedHarm tasks with matched benign controls. On its agent tasks at a 1% false-positive budget, per-step review detects 0.15–0.22 of harmful tasks however many steps it reads, while a random window of four steps detects 0.40–0.50. The window's advantage is significant in eleven of twelve headline comparisons, holds when reviewers read only 18–35% of each workflow, and degrades but survives an agent that rewrites its plan to evade per-step review, which drives per-step detection to near zero. Targeting, the standard prescription in the literature, pays off mainly once evidence is aggregated: at tight budgets the strongest budget-matched policy we test reads the window the monitor finds most suspicious. Monitors should aggregate evidence across steps before deciding, and safety measured on fixed decompositions can overstate safety against agents that choose their own.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.