acceptodds
Under review as a conference paper at ICLR 2027

Safety through Deterrence: Oversight of Colluding Multi-Provider Agents

Abstract

Multi-agent systems can produce artifacts faster than comprehensive trusted verification can check them. Scalable oversight therefore often relies on agent review, but in multi-provider settings reviewers may collude to pass unsafe artifacts. We show that review alone cannot provide a worst-case deterrence guarantee, as a blocking coalition can defeat majority voting, while an agent's pivotal vote flip is invisible to any attribution statistic that is blind to that vote. We therefore place no behavioural assumption on agents and instead target incentives at their providers. Our mechanism reserves a review-independent random audit and penalises the audit-confirmed owner of an unsafe artifact. Any nonzero audit probability then yields a finite penalty that makes planting an unsafe artifact unprofitable, regardless of reviewer behaviour. Reviews are used only for efficiency, with a diversity router allocating the remaining audit budget toward unusually suspicious modules, while a capped bounty incentivises useful reports irrespective of the deterrence guarantee. On vLLM code diffs reviewed by five open-weight models, frontier-authored backdoors evade honest majority voting in 21 of 30 cases. Under collusion, review-dependent architectures eventually reach zero catch probability, whereas the random audit retains its guaranteed floor at every coalition size. With honest reviews, diversity routing raises the catch probability to 90.9%. Allocating a 30% audit budget entirely to the random floor, owner-targeted deterrence requires a penalty of only times the deviation gain, while reserving four of the seven audit slots for the floor and routing the remainder holds honest catch above 90% at a penalty of times. We open-source our code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.