OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
Abstract
AI agents increasingly operate through tools in persistent environments, where each action can alter the state encountered by subsequent actions. As these state changes accumulate over long interaction horizons, safety failures can emerge from sequences of individually benign actions rather than from any single decision. Yet existing safety benchmarks largely evaluate agents on short, static tasks through benchmark-specific interfaces, making it difficult to capture such cumulative risks or compare safety across agent runtimes. To address these limitations, we introduce OpenART, a red-teaming arena based on environment evolution. OpenART constructs over 10K validated, stateful scenarios spanning 50 domains from more than 500K Tools, MCPs, and Skills. These scenarios require a median of 97 tool calls and are projected through adapters onto 15 deployed agents, 5 foundation models, and 8 attack vectors, yielding 75 agent–model configurations. To systematically explore evolving attack surfaces, we further introduce the Evolutionary Markov Hypergraph Attack (EMHA), a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions along hypergraph paths. Throughout this process, the task objective and safety contract remain fixed; only the target-visible environment evolves. Across all configurations, EMHA achieves a pooled strict Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from 1.8–2.7% in simple environments to 17.2–17.6% in the most complex ones. Moreover, target-agent identity explains an additional 7.6% of ASR variation beyond model and capability controls, demonstrating that runtime implementation is itself a consequential factor in agent safety. Together, these results establish OpenART as a scalable foundation for evaluating and red teaming AI agents in complex, stateful, and evolving environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.