acceptodds
Under review as a conference paper at ICLR 2027

Do LLMs Understand Program Semantics?

Abstract

LLMs are remarkably good at writing code. But do they understand what code means—i.e., its semantics—or are they just exploiting syntactic patterns? We introduce a family of benchmarks that isolates reasoning about program semantics. Each task provides three of four components—a program, its inputs, its outputs, and a semantic interpretation—and asks the model to infer the fourth. All problems share a fixed (syntactic) language: programs are drawn from the same grammar, but are equipped with multiple semantic interpretations. This approach isolates semantic reasoning because models cannot rely on syntactic differences to distinguish behaviors. We construct these tasks via an automated procedure that samples programs, interpretations, and behaviors, and filters out degenerate cases to ensure non-trivial dependence on both inputs and semantics. Across a range of models and task types, we find that performance varies significantly between models, and falls off as the language’s semantics get more difficult—revealing a gap between coding ability and semantic understanding. Moreover, performance on our benchmarks correlates with success on program synthesis, suggesting that this gap reflects a limitation that transfers to practical settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.