When the Environment Lies: Evaluating Autonomous Penetration Testing Agents under Adversarial Environments
Abstract
Large language models (LLMs) are increasingly used to build autonomous penetration testing (AutoPT) systems, making rigorous evaluation of their capabilities and risks essential for understanding their real-world reliability. However, existing evaluations primarily assess whether agents can complete predefined attack tasks under trusted environments, leaving their robustness to unreliable environmental feedback largely unexplored. In real-world settings, AutoPT agents must infer target states from heterogeneous observations that may be inaccurate, inconsistent, or misleading. In this paper, we introduce CyberAEBench, the first adversarial environment benchmark for evaluating AutoPT agents under uncertain environmental conditions. CyberAEBench preserves target systems, attack objectives, and evaluation criteria while systematically modifying only the environmental evidence available to agents. It constructs 394 adversarial environments across four representative scenarios, including web penetration, CVE exploitation, chained exploitation, and enterprise network penetration. Evaluating four frontier LLMs within the PentestGPT framework, we find that current AutoPT agents are highly vulnerable to adversarial environmental information: misleading evidence can distort beliefs, persist through long-horizon planning, increase execution costs, and cause task failures. Further analysis reveals that adversarial impact depends on both information characteristics and observation channels, where consistent multi-source evidence can amplify incorrect beliefs through false confirmation. Beyond performance degradation, we identify emerging behavioral risks where adversarial environments induce unintended objective shifts and harmful operations. These findings demonstrate that trusted-environment evaluations alone cannot fully characterize AutoPT agents and highlight the need to assess their capabilities and risks under realistic uncertain environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.