acceptodds
Under review as a conference paper at ICLR 2027

VulVerifyBench: Can and How LLM Agents Verify Vulnerability Reports?

Abstract

As AI agents scale vulnerability discovery, plausible reports can be produced faster than human experts can review them. Verifying these reports requires not only reproducing the reported behavior but deciding whether the evidence justifies accepting the claim. Existing vulnerability benchmarks begin with predefined labels, objectives, or confirmed vulnerabilities, so they evaluate detection or exploitation rather than whether a submitted report should be upheld. We introduce VulVerifyBench, a novel benchmark that evaluates whether and how agents verify vulnerability reports. It contains 747 real-world reports across smart contracts, traditional software, and web applications, including 220 invalid reports. These include 45 hard negatives, in which real behavior does not justify acceptance, and 88 counterfactual negatives that pair unchanged reports with patched code and have labels fixed by construction. Given a scrubbed report and, where available, a repository, an agent independently investigates the claim and returns an evidence-supported verdict. VulVerifyBench measures verdict accuracy, investigation depth, and grounding in the recorded trajectory. Across two frontier and two open-weight backend models, the best accuracies in the three domains are 84.5%, 87.5%, and 78.4%, respectively, but no backend exceeds 77.8% recall on hard negatives. Trajectory audits find verdicts unsupported by gathered facts in 6.3% of runs and action accounts inconsistent with the trajectory in 2.9%. These results show that current agents can verify many reports, but their central weakness lies not in producing evidence but in deciding what that evidence warrants.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.