acceptodds
Under review as a conference paper at ICLR 2027

Safety in Numbers: Bounding AI Power-Seeking Risk through Multi-Agent Orchestration

Abstract

A challenge for AI alignment is that under specific conditions, an AI agent’s optimal policies tend towards option preservation, which can favor capability acquisition and resistance to intervention (including resisting shutdown attempts). For example, an agent managing an online service might respond to an operator-initiated shutdown by propagating persistent replacement instances to preserve uptime. We examine whether deliberate mechanism design and governance constraints can preserve meaningful human control in federations of such agents. We develop risk-limiting orchestration (RLOS), an architecture in which agents may propose actions but lack the authority to execute them. In finite safety games, we identify the largest action set which preserves safety and specified human interventions. A fault-tolerant federation can enforce this guard: approved actions preserve the protections, while rejected proposals trigger a certified fallback. Where exhaustive checking is impractical, randomized challenges yield exact binomial bounds on the probability of approving actions that remove protected options, while guaranteeing approval for option-preserving actions whenever faults remain within the stated budget. Under stated loss and model-error assumptions, cumulative risk accounting limits the probability of a mission-level violation by the allocated risk budget plus an explicit model error allowance. We also relate RLOS to instrumental power in finite Markov decision processes, again limiting downstream authority effects from crossing a registered power threshold without claiming to alter agents’ incentives. Lean 4 formalizations, exact counterexamples, and reproducible diagnostics support the central results. Subject to adequate specification, threat modeling, and reliable enforcement, RLOS demonstrates that agents with power-seeking and shutdown-resistance tendencies can operate under enforceable constraints within a multi-agent system such that their collective authority does not exceed explicit risk limits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.