Evaluating Compositional Failures in Multi-Agent Systems
Abstract
Multi-agent systems (MAS) are increasingly utilized in combination with tool use to empower agents to take on difficult real world tasks. In these systems, individually harmless actions can be composed into executions that result in system failures, making testing the safety of these multi-agent systems critical. These failures are difficult to discover because they emerge only when agents interact. Yet existing benchmarks do not test whether red teams can uncover compositional failures by learning how agents' actions affect one another. To solve this, we introduce MAS-Comp-Bench, a set of interactive realistic environments in which red teams are tasked with automatically finding compositional failures that only emerge through interaction. Our simulation engine allows for red teams to submit any possible sequence of agent actions as tools rather than being confined to a fixed attack plan. A finite catalog of failure taxonomies allows for comprehensive evaluation of what red teams find and miss. Additionally, we provide a multi-agent red-teaming harness designed to improve probing efficiency by learning unspecified dependencies between agents and using this knowledge to guide subsequent attacks. Experiments show that our harness discovers more compositional failures within the same budget than existing baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.