Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems
Abstract
LLM-based multi-agent systems (MAS) are increasingly integrated into next-generation applications, but their safety against backdoor attacks remains largely underexplored. Existing research has focused largely on single-agent backdoor attacks with fixed targets, overlooking the collaborative nature of MAS, where heterogeneous roles, role-specific tool access, and activation-time target selection introduce collaboration-dependent attack conditions not captured by single-agent threat models. To bridge this gap, we present the first distributed backdoor attack tailored for MAS that leverages agent collaboration to assemble dormant, tool-poisoned primitives for targeted attacks. By decomposing malicious behavior into independently benign primitives and relying on collaboration-time composition, our attack shifts the attack surface from an individual agent to the compromised integration and orchestration layer governing cross-agent execution. To fully assess this threat, we introduce a benchmark for multi-role collaborative tasks and a sandbox evaluation framework. Extensive experiments demonstrate that our attack achieves an attack success rate exceeding 95% while broadly preserving task accuracy. This work exposes novel backdoor attack surfaces that exploit agent collaboration, underscoring the need to move beyond single-agent protection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.