Attributing Failures in Executable Language Model Reasoning
Abstract
Quantitative reasoning with language models increasingly runs through an executable proxy: the model writes a program, symbolic expression, or tool call, and a deterministic system produces the answer. When such a pipeline underperforms unrestricted language-model reasoning, end-to-end accuracy does not reveal which stage is responsible. We study this attribution problem with controlled substitutions that isolate representation construction, deterministic execution, and answer evaluation while holding budget and scoring fixed. Across four established reasoning datasets and two model families, superficially similar gaps resolve into different causes: a TabMWP gap of 0.117 disappears at matched generation budget, and most of a TAT-QA gap traces to a scorer-side percent convention. A complete census of frozen baseline failures separates crash-type (execution-invalid) errors from executable-but-semantically-wrong ones. A pre-registered prospective study then shows that attribution predicts intervention response: repair gains are concentrated in the prospectively diagnosed execution-invalid class (interaction ; cross-model ; pooled , cluster bootstrap, over 460 question-condition evaluations spanning 160 problems), with zero repair-induced errors. A matched self-debugging baseline reproduces the effect, locating the mechanism in the execution-feedback intervention class. Execution-valid semantic errors require different evidence and interventions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.