Should This Be an Agent? Policy-Conditioned Autonomy Selection for Enterprise AI
Abstract
Organisations increasingly ask language models what to build for a request: an autonomous agent, a supervised agent, a deterministic rule, a retrieval system, an assistant, or no AI at all. Under a written autonomy policy this design-time choice is a short decision list over facts the request states or omits, yet, to our knowledge, no benchmark measures it under an explicit policy with “no AI” as an admissible answer. We release AUTOSEL-120, 120 enterprise requests with policy-derived reference labels that five AI research scientists independently assessed (three first-round judgments per request; Krippendorff’s α = 0.87; consensus matched 115/120 labels), and Clarify-24, 24 requests each missing one decisive fact. Across six LLMs, 75 of 87 direct errors grant more autonomy than the policy permits (predominant in five of six models), a pattern that survives floor, over-ask, label-artefact and paraphrase controls and is absent from a bag-of-words classifier trained on the same labels; the tested deliberation prompts do not reliably correct it. We then separate reading from deciding: an LLM extracts three-valued decision factors and a strict three-valued compiler applies the policy, so a non-abstaining decision is sound whenever the asserted factors are correct, a missing decisive factor becomes a named deferral, and a policy change over the same factors is a recompile. The policy also compiles into model-certified training data: defer-aware 8B extractors make 3/4/3 wrong decisions on 120 requests across three seeds while deferring 17/20/20 times (coverage 83–86%, selective accuracy 96–97%), but they defer on 17/17/16 of the 20 no-AI requests (D0 recall 1/20) and detect only 5/5/6 of the 24 under-specified requests (naming the missing fact in 5/4/5). Benchmark, judgments and code are provided as anonymous supplementary material; coverage and reliable detection of under-specified inputs remain the open limitations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.