The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark
Abstract
AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequential to cybersecurity, including malware, firmware, and proprietary applications, is available only as binaries. Analyzing such software requires reverse engineering (RE): recovering program semantics before the analysis can be meaningfully performed. However, evaluating agentic RE poses a fundamental challenge: realistic and useful benchmark instances must (1) be absent from LLMs’ training data to prevent models from taking shortcuts by memorization, and (2) represent the scale and anti-analysis protections of real-world binaries. To faithfully evaluate the RE capability of frontier models, we introduce SRE-Bench, the first realistic, contamination-free RE benchmark. Built entirely from scratch by RE experts with over 5,000 expert hours, SRE-Bench comprises 19 private, real-world-scale programs averaging 16.9K lines of code. We further developed 44 in-house anti-analysis mechanisms, yielding 262 binary instances and 1,572 deterministically graded tasks. We evaluated 13 agentic settings across 11 models: eight in public-facing settings and five in internal unconstrained settings with cyber safeguards disabled and no budget cap. Our results show that realistic RE remains challenging for frontier agents: GPT-5.6-Sol and Claude-Fable-5.1, despite their strong source-code cybersecurity capabilities, fully solve only 31.5% and 26.9% of graded instances, respectively. This result suggests that success in source-code security does not necessarily translate into effective binary analysis. Without a budget cap and safety guard, GPT-6-Astra achieves a near-perfect pass@4 score; however, reliably identifying the correct candidate still poses a significant challenge. Our analysis further reveals that agents are relatively insensitive to compiler optimization and static linking, and controlled ablations confirm that both contamination control and realistic scale are essential to understand agents’ RE capability. Together, these findings highlight RE as a distinct frontier for agentic cybersecurity and establish SRE-Bench as a rigorous testbed for measuring progress.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.