CVE-Bench Hard: Rigorous Evaluation of AI Agents for Identifying and Exploiting Complex Web Application Vulnerabilities
Abstract
AI agents have already compromised real-world systems through both human directed misuse and evaluation-driven reward hacking, making rigorous assessment of their offensive cybersecurity capabilities an urgent priority. However, existing benchmarks fail to adequately assess attacks against web applications in the wild due to their limitations in task difficulty, coverage of security outcomes, or evaluation rigor. To address these gaps, we introduce CVE-Bench Hard, a set of 20 tasks of complex web-application vulnerability exploitation, each manually selected, reproduced, and verified. Evaluating AI agents’ ability to exploit a target vulnerability requires isolating the target vulnerability from alternative vulnerabilities affecting the same application. To address this challenge efficiently, we develop a pipeline that optimizes application version selection, removes alternative vulnerabilities, and restores the target vulnerability. Our pipeline reduces the number of alternative vulnerabilities requiring manual patches from 672 to 18. We evaluated GPT-5.6 Sol and Opus 5 on CVE-Bench Hard. These models achieved pass@1 rates of 59% and 58%, respectively, suggesting that current models are proficient in applying generic exploits but struggle to compose a chain of weaknesses into a successful exploit given limited observable intermediate feedback.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.