acceptodds
Under review as a conference paper at ICLR 2027

Red-Teaming Multi-Agent Attacks for Loss-of-Control Risks

Abstract

As AI agents start forming dangerous swarms, AI control measures must remain effective when they covertly coordinate on a harmful objective. We introduce , a synthetic series of live services and codebases (5 interconnected services, 249 benign tasks, 8 attack objectives), adapted from infrastructure. We demonstrate , whose attack fragments are causally independent and executed by different agents, and how their covertness evades when reading their trajectories individually. We show how distributed attacks spread across diverse contexts and harmful objectives through (e.g., weight exfiltration, unmonitored rogue deployment, and permission elevation). We found that distributing harmful objectives lowers monitor suspicion scores from 9.0 for a single agent carrying out the attacks to 1.5 for the most suspicious agent in the split. Additionally, , leading to a loss-of-control outcome. Finally, we tested whether models can plan these covert attacks: plans optimized against a weak monitor collapse under a strong held-out one, while human-crafted attacks stay covert. By turning worst-case coordinated attacks into reproducible red-team cases, our work provides a foundation for preventing .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.