acceptodds
Under review as a conference paper at ICLR 2027

Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors

Abstract

LLM-based vulnerability detectors now review code automatically before it is merged. We ask whether an attacker can make such a detector miss a vulnerability that it would otherwise flag, using only edits that keep the code compilable and leave the vulnerability in place, such as renaming a variable or adding a comment, an inactive preprocessor directive, or a branch that never runs. We restrict universal adversarial-string optimization to these four carrier families and evaluate six detectors on 5,000 vulnerable C/C++ functions. Five of the six detectors lose 87% or more of the vulnerabilities they detect to at least one of the five resulting attack variants. Strings optimized once on a 14B open-weight surrogate evade 87.09% of the vulnerabilities that GPT-4o detects, and optimizing on the detector itself evades up to 99.88%. No single variant is strongest on every detector, and clean recall does not predict how many detected vulnerabilities survive all five.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.