acceptodds
Under review as a conference paper at ICLR 2027

Learning a Symbolic Decompiler

Abstract

Decompilation reverses the transformations of compilers. These transformations are deterministic but non-bijective, often discarding names, types, and structure from high-level languages. Traditional decompilers infer plausible high-level information through heuristic analyses that require expert labor to write and maintain. Recent work instead fine-tunes large language models to generate source code directly or uses LLM agents to iteratively refine the output of existing decompilers. Both approaches make decompilation knowledge hard to accumulate and inspect: neural models encode it implicitly in weights, where individual behaviors are hard to trace or repair; agents typically reason about each program independently, leaving no residue for future tasks. Symbolic rules offer complementary properties: they are inspectable, traceable, precisely repairable, and reusable across programs, but have traditionally been difficult to learn autonomously. To address this, we present RuleCrafter, an agentic framework that decomposes decompilation into substeps and learns Datalog rules for each. Without LLM calls at evaluation, the best learned rule base matches or exceeds traditional decompilers in re-execution, and rule bases learned on one corpus apply to the others without modification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.