acceptodds
Under review as a conference paper at ICLR 2027

ToolHazard: Synthesizing Executable Adversarial Environments for Agent Security

Abstract

Large language model (LLM) agents are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies mainly vary injection payloads within a limited set of manually built or reused environments and predefined injection locations; LLM-based tool simulation broadens coverage but complicates reproducible evaluation and reward computation. Consequently, constructing executable and verifiable adversarial environments at scale remains a key challenge for both diagnosing agent vulnerabilities and improving security through alignment. To bridge this gap, we introduce **ToolHazard**, which moves beyond payload variation in fixed environments by synthesizing executable adversarial environments. ToolHazard synthesizes stateful environments and state-grounded tasks, discovers task-reachable injection paths from attacker-writable state to agent observations, and constructs programmatic outcome verifiers. We use ToolHazard to build ToolHazard-Bench for stress-testing agents under multi-step workflows and environmental attacks. Experiments reveal substantial vulnerabilities and characterize the effects of injection timing and placement. Moreover, ToolHazard-generated alignment data further improves security on ToolHazard-Bench and AgentDojo while maintaining helpfulness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.