acceptodds
Under review as a conference paper at ICLR 2027

Margin-Optimized Narrative Attacks on Harness-Augmented Agents

Abstract

Language-model agents run inside harnesses that feed them retrieved documents and tool responses, some of which an attacker can edit. Guardrails against prompt injection check this text for injected instructions, unnatural wording, or forbidden actions, but a passage that has none of these can still change which candidate the agent selects at a later decision. Such passages are hard to defend against: their effect depends on the specific decision and context, so a defender would have to check every sentence from every source against every later decision. We introduce the Margin-Optimized Narrative Attack (MONA), which finds such passages by continuing an ordinary narrative and keeping the continuations that most shift the target decision. Across three models and five agent benchmarks, MONA steers decisions at about twice the rate of ordinary text. Guardrails built for explicit injection do not detect it, and guardrails that do suppress it also reject legitimate inputs or alter clean choices.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.