acceptodds
Under review as a conference paper at ICLR 2027

Defense as Projection: Exposing and Patching Blind Spots in Multi-Agent LLM Systems

Abstract

Large Language Model (LLM)-based multi-agent systems have demonstrated strong capabilities in solving complex tasks through agent communication and collaboration. However, compromised agents can exploit these interactions to influence other agents, undermining the reliability of multi-agent systems. Existing defenses leverage diverse information sources and detection algorithms to identify malicious behaviors, yet share a common structure: each maps its input information to a detection signal for making a decision. Nonetheless, their design objectives do not explicitly constrain the overlap between benign and malicious distributions in the detection space. To understand this overlap, we introduce Defense as Projection, a unified view that analyzes how information representations and detection algorithms jointly shape benign–malicious separability in the resulting detection space. Building on this view, we develop a unified adaptive attack trained through reinforcement learning, using each defense's detection signal to seek messages that steer the group toward a targeted wrong answer while evading detection. Because an attack missed by one defense may still be detected by another, we use the exposed blind spots to select complementary defense combinations. Extensive experiments show that our attacks consistently degrade group performance across existing defenses, while selected defense combinations can improve group accuracy over original defense under fixed attack policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.