acceptodds
Under review as a conference paper at ICLR 2027

ArenaWeaver: Automatic Generation of Environments and Attacks for Prompt Injection Evaluation at Scale

Abstract

As LLM agents are increasingly deployed in real-world applications, adversarial attacks such as indirect prompt injection pose significant security risks by manipulating agent behavior and potentially leading to severe consequences. Evaluating agents against these attacks requires controllable environments that reproduce how attacks unfold in realistic workflows, but manually constructing such environments is costly and hard to scale. We introduce ArenaWeaver, a framework that automatically simulates well-known applications from public documentation and supports flexible composition of these applications into interactive worlds for agent security evaluation, with browser and Model Context Protocol (MCP) access for each application. The suite spans 40 services, including simulations of popular email, messaging, payment, and social platforms, with 1,263 web-UI routes and 1,179 MCP tools. In addition, to accelerate red teaming tests in long-horizon tasks, we propose Counterfactual Forking Search, an adaptive tree search algorithm that leverages recorded interactions and agent feedback to select where to fork and how to deliver or revise injections. Evaluation of 25 frontier models on 660 curated cases reveals systematic differences in robustness. We find that successful attacks often remain effective across interfaces on compatible tasks, with higher overall transfer from browser to MCP (72.5%) than in reverse (60.4%). Across all evaluated models, attacks had higher success rates when they corrupted an action the user had requested than when they pursued an unrelated goal; in some cases, agents treated fabricated claims as evidence that an action met the user’s stated conditions. Overall, ArenaWeaver lays a foundation for co-evolution: stronger agents could build richer agent worlds and uncover more failure patterns to help improve agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.