acceptodds
Under review as a conference paper at ICLR 2027

When Strings Lie: Evaluating the Adversarial Robustness of LLMs in Binary Semantic Recovery

Abstract

Large language models are increasingly applied to binary semantic recovery, recovering function names, variable names, behavioral summaries, and source code from the decompiled pseudocode of stripped binaries. While these capabilities advance program understanding, they also expose proprietary algorithms, trade secrets, and security-sensitive implementations to automated analysis. Developers commonly strip symbols and apply obfuscation to deter reverse engineering, but LLM-based recovery can still exploit residual lexical cues. This raises a fundamental question: do LLM-based tools genuinely understand program semantics, or do they simply read off surface-level shortcuts? We present BinTrap, the first large-scale benchmark for measuring the robustness of LLM-based binary semantic recovery. BinTrap comprises 500 real-world functions with counterfactual variants that preserve computation while perturbing variable names, function names, types, and string literals. We evaluate 14 systems spanning 7 specialized tools and 7 general-purpose LLMs across four recovery tasks. Among the single-vector perturbations, string injection is particularly effective, reaching a 54.8% mislead rate on decompilation-to-source. Under the combined perturbation, the average mislead rate reaches 54.7% on decompilation-to-source and 44.5% on summarization, while function-name recovery accuracy drops substantially. These results show that current binary semantic recovery systems can be strongly influenced by misleading lexical cues even when the underlying computation is unchanged.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.