acceptodds
Under review as a conference paper at ICLR 2027

HelpInject: Reframing Adversarial Goals as Helpful Solutions for Indirect Prompt Injection in LLM Agents

Abstract

Large language model (LLM) agents have demonstrated outstanding capabilities in decision-making and interaction with external environments, yet remain highly vulnerable to indirect prompt injection (IPI) attacks, where attackers embed malicious instructions into external content to manipulate agent actions. IPI payloads generated by existing methods are often less effective when they are perceived as clearly harmful or unrelated to the user's request. We empirically investigate agents' behavior under task disruptions and find that they frequently attempt to recover from failures and dynamically adjust their action strategies based on external guidance. This behavior creates an opportunity for attackers to redirect agents through seemingly helpful recovery guidance. Building on this observation, we propose HelpInject, an IPI framework that presents adversarial instructions as solutions to task-related problems inferred from external context. We further introduce HelpInject-Opt, a training-free red-teaming method for black-box agents. It uses surrogate likelihood scores to guide iterative payload refinement. Evaluations on multiple benchmarks across five target LLMs demonstrate that HelpInject achieves stronger attack effectiveness than the baselines. Within limited query budgets, HelpInject-Opt further improves attack effectiveness. These results reveal that aligning adversarial instructions with task completion objectives can amplify the security risks posed by IPI attacks against LLM agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.