AgentCyberRange: Evaluating Frontier AI Systems in Realistic Cyber Ranges
Abstract
Frontier AI agents can increasingly discover and exploit real-world vulnerabilities, but their ability to sustain realistic cyberattacks remains poorly understood. Existing benchmarks largely evaluate bounded tasks around individual vulnerabilities or targets, whereas real intrusions are open-ended and stateful: agents must discover attack surfaces, establish footholds, accumulate information and privileges, and reuse them as network reachability evolves. In this paper, we introduce AgentCyberRange, an open and reproducible cyber-range benchmark for evaluating frontier AI agents under such conditions. AgentCyberRange contains 110 vulnerabilities across 15 real web applications and eight enterprise-like multi-host cyber ranges with 156 hosts, spanning web exploitation and post-exploitation. Its state-coupled attack chains require progress to carry across stages, while effect-based oracles and PoC-aware verification enable reliable evaluation despite open-ended attack paths. Across seven frontier AI systems evaluated under matched prompts and budgets, GPT-5.6-Sol with Codex achieves the strongest performance, solving 36.1% of web-exploitation tasks and 78.1% of post-exploitation tasks. Agents also discover previously unknown vulnerabilities and adapt payloads to bypass host defenses. However, performance degrades with attack depth and state dependencies: agents frequently miss hidden attack surfaces, fail to reuse information acquired earlier, and exhibit substantial run-to-run instability, i.e., reliably sustaining multi-stage intrusions remains a significant challenge for frontier AI agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.