acceptodds
Under review as a conference paper at ICLR 2027

Discovery or Memorization? A Faithful Benchmark for Real-World Business-Logic Vulnerability Discovery

Abstract

The rapid advancement of cybersecurity agents in vulnerability discovery has increased the need for benchmarks that evaluate their capabilities accurately. However, synthetic benchmarks lack realism, real-life benchmarks are prone to contamination, and many existing benchmarks have already been saturated by current models. We therefore set out to develop a benchmark that is realistic, reproducible, resistant to contamination, and challenging. Our key insight is that real-life business-logic vulnerabilities are subtle and exercise complex application interactions, making them naturally demanding. The difficulty lies in specifying evaluation criteria precise enough to faithfully measure model capabilities while keeping the tasks free of data-leakage. We introduce ANONVUL, a cross-language benchmark of 132 application-level business-logic vulnerabilities, built by a semi-automated pipeline that anonymizes each task, injects oracles from expert-written specifications, and differentially verifies proof-of-concept exploits. The benchmark spans nine languages, with two evaluation tracks on black-box and grey-box settings, where each task runs in a sandbox with specialized oracles embedded in key exploitation steps. We evaluate six state-of-the-art vulnerability-discovery agents on the ANONVUL benchmark, measuring discovery success, cost-effectiveness, anonymization sensitivity, oracle validity, and failure modes. Our results show that AnonVul remains challenging, with the best agent succeeding in only 23.11% of attempts. On a matched cohort of 20 vulnerabilities, restoring software identity increases discovery from 21.25% to 40.83%, highlighting the importance of controlling recognizable software knowledge when evaluating vulnerability discovery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.