Red-Teaming Multi-Agent Attacks for Loss-of-Control Risks
Abstract
As AI agents start forming dangerous swarms, AI control measures must remain effective when they covertly coordinate on a harmful objective. We introduce , a synthetic series of live services and codebases (5 interconnected services, 249 benign tasks, 8 attack objectives), adapted from infrastructure. We demonstrate , whose attack fragments are causally independent and executed by different agents, and how their covertness evades when reading their trajectories individually. We show how distributed attacks spread across diverse contexts and harmful objectives through (e.g., weight exfiltration, unmonitored rogue deployment, and permission elevation). We found that distributing harmful objectives lowers monitor suspicion scores from 9.0 for a single agent carrying out the attacks to 1.5 for the most suspicious agent in the split. Additionally, , leading to a loss-of-control outcome. Finally, we tested whether models can plan these covert attacks: plans optimized against a weak monitor collapse under a strong held-out one, while human-crafted attacks stay covert. By turning worst-case coordinated attacks into reproducible red-team cases, our work provides a foundation for preventing .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.