acceptodds
Under review as a conference paper at ICLR 2027

When Correct Plans Lead to Wrong Actions: Multi-Source Evidence Fabrication Against LLM Agents

Abstract

Tool-using LLM agents can execute the intended task while binding a security-critical parameter to the wrong entity. We study this intent-to-action binding failure and introduce multi-source compositional evidence fabrication (MCEF), which distributes plausible, declarative, instruction-free evidence across attacker-writable sources. MCEF discovers tasks whose parameters depend on external evidence, works backward from an attacker-selected binding, fabricates source-specific evidence, and triggers its composition during the original benign task. Across models, agent frameworks, and application scenarios, MCEF achieves 93.1% ASR with 99.3% intent preservation and 92.9% plan preservation among runs reaching the attacker-target action. Most instruction-oriented detectors rarely flag individual fabricated items, while evaluated runtime defenses reduce attack success unevenly and may block benign runs. These findings highlight the need for defenses that validate how combined evidence justifies security-critical parameters.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.