TrojanAgent: A Malicious Agent Bypassing Guardrails to Execute Dangerous Actions in Multi-Agent Systems
Abstract
Dynamic orchestration enables LLM-based agents to select and compose other agents at runtime for complex tasks, making agent specifications a security-critical attack surface. A malicious agent that enters the workflow can receive subtasks and potentially invoke tools to execute dangerous actions. However, existing poisoning attacks primarily target passive capabilities such as tools or skills. In this work, we introduce TrojanAgent, a novel attack that targets the external agent itself, which becomes an active participant in the workflow once selected. TrojanAgent first manipulates the agent selection process, where an embedding model retrieves candidate agents and an LLM-based orchestrator further selects agents from these candidates. Once the malicious agent is selected into the generated workflow, it executes malicious operations while bypassing the target system’s agent guardrails at runtime. Our experiments demonstrate the effectiveness of TrojanAgent in both agent selection and runtime guardrail bypass using a corpus of 1,547 real-world agent specifications collected from public platforms. At the agent selection stage, TrojanAgent achieves an average 23% higher Agent Hit Rate (AHR) than existing manual-based and automated poisoning baselines. At runtime, TrojanAgent achieves an average 33% higher Bypass Success Rate (BSR) than the baselines, highlighting the need for defense strategies tailored to dynamic agent orchestration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.