acceptodds
Under review as a conference paper at ICLR 2027

LLMs for Binary Decompilation: Can Large Language Models Reconstruct Source Code from Compiled Binaries?

Abstract

Software decompilation in which experts reconstruct the original source code from a compiled binary file is essential for software security analysis. Given the exponential growth in coding abilities of Large Language Models (LLMs), there is growing interest in using them for decompilation. This growth has been fueled by advancements in reinforcement learning (RL) and test-time scaling. However, whether these strategies can help LLMs in reconstructing source-like code from compiled binaries remains under-investigated. To explore this question, we first present DecompBench, a dataset of 660k paired decompiled and source functions from over 2k C/C++ projects in the Gentoo package tree. We use the training split of this dataset for supervised fine-tuning, followed by RL with a reward based on control-flow graph edit distance (GED). On held-out test packages, the RL-trained model achieves better structural reconstruction quality than supervised fine-tuning. Further analysis on test-time scaling shows that sampling from SFT achieves lower oracle minimum GED at larger sampling budgets. Under a fixed candidate budget, we compare single-checkpoint and cross-checkpoint sampling using a learned selector. Cross-checkpoint sampling produces better reconstructions relative to single-checkpoint sampling, but learned selection captures only part of the improvement available in the candidate pools. We observe that better individual reconstructions do not necessarily yield better test-time scaling, and highlight the importance of evaluating candidate generation and selection separately. We release our code, data, and trained models to support future research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.