A Wrong Turn in Addition: Generation, Routing, and Correction of Arithmetic Errors in LLMs
Abstract
Despite rapid progress in mathematical reasoning, large language models remain surprisingly brittle on elementary arithmetic. We ask how a local arithmetic error becomes a persistent autoregressive trajectory, and whether evidence of that error survives after the model has committed to it. Using matched digit-preserved carry counterfactuals and causal interventions, we decompose multi-operand addition errors into generation, read-back routing, and downstream compatibility. Natural off-by-one carry errors are characterized by weakened carry-sensitive influence, while the emitted digit subsequently selects between local lower- and higher-carry continuation routes through a localized query–key interaction. Crucially, error propagation and error representation coexist: the model can continue along a wrong route while downstream states remain sensitive to whether the read-back digit is arithmetically compatible with its context. This delayed compatibility signal is visible in next-token behavior, entropy, and representation geometry, and enables a training-free mechanism-guided correction method that substantially improves exact-answer accuracy and outperforms most training-required baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.