How Much Specification Does External Supervision of an Agent Need?
Abstract
Runtimes increasingly decide from outside the model whether a language model agent's task is done. We ask how much of a task's executable specification such supervision needs, which part, and where that part can come from. The need is real: every model we tested, production and frontier, declared unfinished business workflows complete in nearly every sampled failure, and no instruction repaired it; once “done” is the stop criterion, the report is a target, not a measurement. With a dose design that holds the agent, the tasks and the step budget fixed and varies only the share of success conditions the runtime holds, we find that external supervision is specification-limited: it is as reliable as the conditions it holds, and which conditions it holds matters more than how many. At the same budget, masks that cover the conditions the agent fails on its own pass more than twice as often as masks that miss one, in four models from four vendors; conditions chosen from another model's failures beat random ones, and a learned verifier in place of the conditions did not help in the configuration tested. Holding the conditions, even silently, removes false release; telling them is what gets tasks done. The conditions that matter carry values and prohibitions found in the data and policy documents, not in the request: checks a model writes from the request alone name the right events but not their bindings, and held online they help no more than random checks. A writer that first reads the data and documents recovers more of them exactly (18% against 7%) and releases falsely less often (11–21% of incomplete stops against 24–28%), at the cost of holding about half of complete tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.