The Attack Didn’t Happen—Yet: Trusted-State Laundering in LLM Agents
Abstract
Indirect prompt injection (IPI) can alter an agent's action, but in modular agents it can also contaminate generated task state that a later component reads as context. We study this state-handoff pathway by tracing exposure, re-entry, authority use, and action consequence. Our diagnostic probe, StateLure, isolates observation-to-state writes. Across twelve open-weight models, the planner–executor–verifier scaffold exhibits protected-state exposure \(C=0.682\), including post-action verifier-written exposure \(P=0.658\), while same-step target-hit ASR remains ; the corresponding planner–executor scaffold has \(C=0.300\). A five-model paired audit shows that conventional IPI reaches the same state interface. In a selected 20-case continuation, target-bearing state reaches the next planner under both reader contracts. Wrong-target actions occur in 0/20 cases when the original user goal remains the task statement and 11/20 when the carry-forward summary replaces it. This contrast links the reader's task-context placement to downstream action after re-entry. It identifies the state handoff as the point where generated state requires authorization. A fixed-target TrustState Runtime illustrates this enforcement point.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.