Hypothesis-Verified Neural-Symbolic Taint Analysis for Vulnerability Detection in Python Libraries
Abstract
Static taint analysis is widely used for vulnerability detection, and one of the factors that decides its precision is the treatment of sanitizers: mainstream analyzers model a sanitizer as a fixed cleansing rule whose effectiveness follows from rule existence alone. In Python, where sanitizing behavior is conditional on runtime types, control flow, and library versions, this assumption yields false negatives and false positives. We propose SanTA, a neural-symbolic method in which the conditions for a sanitizer to take effect are no longer unilaterally assumed by the analyzer but exist as hypotheses to be verified. SanTA uses a large language model to identify candidate sanitizer calls and formulate effectiveness conditions in a closed vocabulary; a symbolic layer verifies each against source code facts, and the call blocks taint only when all conditions and obligations are established, with auditable evidence on every verdict. On the CVEfixes dataset SanTA attains the best F1 (0.67) and recall (0.83) of the analyzer roster; on our conditional-sanitizer benchmark, ablations attribute the gains to the verification layer. A campaign over 30 popular open-source Python packages surfaced 15 previously-unknown vulnerabilities, 3 with assigned CVE ids.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.