Representation Sensitivity in Symbolic Execution by Large Language Models
Abstract
Large language models can execute symbolic procedures yet remain sensitive to how those procedures are represented. We study this gap using 256 deterministic table-execution programs with exact gold states. In the primary intervention, programs, table values, initial states, and decoding settings are fixed; only the names of row and column roles change. On gpt-oss-120b, replacing overlapping A/B role labels with disjoint P/Q labels raises accuracy from 82.4–86.7% to 98.0–98.8% across three sampling seeds. The direction persists under reversed condition order and across tested Qwen configurations. A subsequent 2×2 experiment crosses A/B versus W/X role labels with A–D versus W–Z value alphabets, preserving computations under a consistent symbol renaming. Across two sampling seeds, the preferred label pair reverses with the value alphabet: overlapping labels produce more errors within both alphabets. Mean within-alphabet overlap penalties are 7.81 and 10.94 percentage points, with paired bootstrap 95% intervals of [3.91, 11.72] and [7.03, 14.84]. All primary and crossed responses terminate naturally and are syntactically valid. A selected 12-root diagnostic cohort localizes the first visible error in ten failing traces to table-cell lookups. These results establish representation sensitivity and a replicated label–alphabet interaction in the tested setting, while leaving internal mechanisms unresolved. They motivate evaluating representation robustness alongside nominal symbolic-execution accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.