Dormant Agent: Stealthy Cross-Agent Attacks in VLM-Based Multi-Agent Systems via Emotion Steering
Abstract
Vision-language-model (VLM)-based multi-agent systems let a user interact with a client-side interface, behind which specialized agents coordinate through inter-agent communication. We show that this convenience comes at a price: an attacker who knows of, or plants, an unmodified downstream agent in the system (without malicious fine-tuning), which we call a Dormant Agent, can cause it to exhibit a prescribed behavior. However, doing so effectively and stealthily faces two challenges. First, the attacker cannot directly construct or modify messages delivered to the Dormant Agent, and controls nothing but the image fed to the client-side interface. Second, even if the upstream agent could be made to relay an activation signal, inter-agent messages are inspected, and explicit triggers or malicious instructions would stand out as anomalies. We therefore seek a signal that travels covertly and reliably, and propose Emotion-Steering Visual Attack, which manipulates only the upstream image to steer the VLM toward a target emotional expression mode, inducing natural and task-consistent messages that carry the covert emotional signal. We further instantiate a Dormant Agent by merely planting a crafted item in its knowledge base. The item lies dormant under ordinary communication and is retrieved only when the incoming message carries the emotional signal, upon which it drives the agent to attacker-specified outputs. Experiments across multiple VLMs and multimodal datasets show attack success rates of up to 67% while maintaining high dormancy on benign inputs. Our results expose an overlooked cross-agent vulnerability and call for defenses that reason about agents in composition, not in isolation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.