ATLAS: Atomized Topology-aware LLM Attack Search for Adaptive Attacks on MAS
Abstract
Inter-agent communication in LLM-based Multi-agent Systems (MAS) creates new attack surfaces for indirect prompt injection. Existing attacks optimize adversarial inputs or exploit communication structures, but offer limited insight into how adversarial influence takes effect during these interactions. To address this gap, we analyze execution traces across representative security benchmarks and identify node-level attack conversion and effective semantic propagation as two key factors consistently associated with attack success. Building on these findings, we propose ATLAS, an Atomized Topology-aware LLM Attack Search framework for adaptive multi-point attacks. ATLAS decomposes an injection into coordinated semantic atoms, jointly searches their injection locations and propagation paths over the communication topology, and preserves the complete attack intent through terminal semantic closure. Experiments across three security benchmarks and ten victim models show that ATLAS substantially improves attack effectiveness while generally preserving benign-task performance and injection stealthiness. These results demonstrate how trace-level analysis of adversarial influence can guide the design of more effective attacks on multi-agent systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.