DF-Fuzz: Dependency-Guided Context Retrieval for LLM-Assisted Language Processor Fuzzing
Abstract
Language processor fuzzing assisted by large language models (LLMs) has shown substantial promise in synthesizing test programs that satisfy high-level syntactic and semantic specifications to penetrate deep code validations. However, existing LLM-driven constraint-solving fuzzers face a severe semantic scarcity bottleneck when testing large-scale compiler systems. Conventional physical function slicing retains only the enclosing function containing the target bottleneck basic block, severing critical semantic linkages to external type declarations, preprocessor macros, and class hierarchies. Consequently, LLM-generated test programs suffer from low compilation pass rates. Meanwhile, interactive retrieval designs incur prohibitive round-trip latency and token overhead, frequently triggering format hallucinations and state-machine deadlocks. To overcome these limitations, we present DF-Fuzz, a dependency-guided context retrieval framework for language processor fuzzing. DF-Fuzz introduces the Dual-Flow Hybrid Dependency Graph (DF-HDG) to unify intra-procedural control constraints and inter-procedural, cross-file data dependencies within an offline-constructed representation. Operating via a dual-track architecture, DF-Fuzz performs reverse dominator tree slicing to isolate the minimal condition chains governing branch reachability, while concurrently employing inter-procedural symbol resolution with namespace wrapping to reconstruct self-contained scopes. To mitigate context bloat, DF-Fuzz integrates Jaccard-similarity-based semantic pruning with an API co-occurrence few-shot retrieval mechanism. Furthermore, DF-Fuzz incorporates an adaptive feedback-driven self-healing loop that translates local compiler diagnostic errors into dynamic exemption weights, repairing omitted context in subsequent iterations at zero additional LLM interaction cost. We evaluated DF-Fuzz across eight real-world language processors spanning six programming languages, including Clang++, Flang, Hermes, and Lua. In 24-hour comparative evaluations, DF-Fuzz achieved 100,153 edge coverage on Clang++, outperforming AFL++ by 78.8% and the state-of-the-art LLM baseline HLPFuzz by 5.7%, while reducing API operational costs by 60.3% and increasing execution throughput by 26.9%. In extended testing across all targets, DF-Fuzz discovered 45 previously unknown vulnerabilities, including 10 heap use-after-free and 8 heap buffer overflow vulnerabilities. Our code and artifacts are publicly available at https://osf.io/dkqz4/overview?view_only=d75600e5f99f41a6a849bbc36515e919.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.