acceptodds
Under review as a conference paper at ICLR 2027

Where You Inject Matters: The Trust Region Effect in LLM Agents

Abstract

LLM agents are increasingly adopted in long-horizon tasks and interact with diverse external environments. Despite their growing capabilities, these agents remain vulnerable to real-world attacks such as indirect prompt injections. Indirect prompt injections manipulate LLM agents through malicious instructions embedded in their external environments. While existing risk assessment work has primarily focused on the injection content, the effect of where an injection appears has not been systematically characterized. Studying injection placement requires agents to operate over longer-horizon tasks and interact with diverse environmental surfaces. Thus, we build three diverse agentic testbeds spanning tool use, coding, and daily workflow domains, where long-horizon interactions expose agents to diverse environmental surfaces. Concretely, holding the injection content fixed, we find that agent susceptibility varies across injection placements. We refer to this phenomenon as the trust region effect: agents implicitly assign different levels of trust to different regions of their environments. For example, an agent powered by GPT-5.4 follows an unsafe injection instruction placed in a repository documentation file (i.e., a README or TODO file) 2.6× more often on average than the same instruction placed in a code or data file. Building upon this observation, we further develop learning-based attackers that generate injection content and select better injection placements for advanced risk assessment. Our approach demonstrates that both where an injection is placed and what content it contains can affect the strength of indirect prompt injection attacks. To the best of our knowledge, this work provides the first systematic characterization for the effectiveness of different injection placements, spanning both coarse-grained environmental surfaces and fine-grained positions. The uneven trust that agents implicitly place across different regions of their environments provides important insights for future attack and defense design for agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.