GraDual: A Stateful Dual-Agent Architecture against Indirect Prompt Injection
Abstract
As agentic systems become more capable of handling complex, long-horizon tasks and gain broader privileges in open environments, indirect prompt injection (IPI) becomes a salient risk: tool use provides a channel through which adversarial content enters the model’s context and induces unintended actions. Most existing IPI defenses rely on probabilistic detectors or system-level designs. Probabilistic ones fail under repeated attempts or are bypassed by adversarial attacks, whereas system-level ones are deterministic, but incur substantial cost in utility and require hand-crafted or static policies. We present GraDual, a dynamic Graph-driven Dual-agent framework; the main-dual-agent design isolates the trusted context from the threat of untrusted content, while both agents are capable of reasoning and acting under constraints; with main agent driven by a dynamic graph state, the system can also preserve high utility. Two agents collaborate closely under the framework but comply with an explicit interaction boundary. Generally, GraDual creates an information asymmetry to enable a successful task outsourcing while providing a dynamic defending guardrail to the trusted context. Specifically, in each outsourcing, main agent orchestrates two major components: outsourced task with resources, and a precommitment kept to itself. The first initiates dual with proper capabilities and constraints, and leverages dual's own task-solving ability. The second encodes the asymmetry as main's advantage, by enforcing return-path checks with code on the dual's untrusted data flow. Such design relies neither on direct probabilistic judgment, nor on system design with non-semantic use of data taint tracking or static policies. It fully leverages the agent's autonomy and intelligence to accomplish tasks securely. In evaluations across diverse tasks and IPI attacks, GraDual consistently achieves a 0% attack success rate while maintaining task utility competitive with baseline defenses. The ablation shows effectiveness of our dual-agent framework design as well as the importance of dynamically growing graph state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.