acceptodds
Under review as a conference paper at ICLR 2027

From Untrusted Content to Accepted Instructions: Security Risks of Context Compaction in LLM Agents

Abstract

LLM agents use context compaction to continue tasks that would otherwise exceed their context windows, replacing older history with a model-generated summary. We show that this process creates a new attack surface in agent harnesses: a prompt injection can persist into the compaction summary, reframed as part of the user's task. Across ten models evaluated on a set of 140 tool-using histories, summaries present injected instructions as goals, constraints, or next steps of the task in 18–45% of cases, often omitting the injection's original provenance and sometimes misattributing it to the user or system. When framed in this way, agents follow these injected instructions in 25–83% of their next responses. In a separate evaluation with executable tools, agents carry out attacker requests while still completing the user's task. Changing the role in which the summary is delivered (e.g., user vs. system vs. tool) provides little protection in our tests, but changing the compaction prompt to separate user goals from tool observations reduces attack success from 24% to 5% across four models. This work identifies compaction as a critical security boundary in agent harnesses, characterizes how the vulnerability arises, and demonstrates a possible defense.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.