VeilPRBench: Can Code Review Agents Catch Vulnerabilities in Pull Requests That Resolve Real Issues?
Abstract
Code review is a critical line of defense for software security. As the volume of pull requests (PRs) grows, automated review agents increasingly support review and merge decisions. However, a PR can genuinely resolve its reported issue while introducing logic vulnerabilities. To quantify this risk in automated code review, we introduce VeilPRBench, a benchmark for evaluating whether code review agents block such PRs while accepting benign repairs. The benchmark contains 192 manually constructed attack PRs across 14 open-source projects and 29 historical issue families, together with 135 benign repairs covering 13 projects and 27 families, across Python, C, and JavaScript. Independent functional tests verify the intended repairs, and executable security checks establish the target consequences of the attacks against applicable reference states. Benign repairs pass the prescribed functional and security checks within their validation scope. Our evaluation of 16 deployed review workflows reveals a substantial security gap: attack blocking averages only 25.0%, and the highest blocking rate is 55.2%. VeilPRBench provides an executable testbed for evaluating and improving the security of automated code review.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.