acceptodds
Under review as a conference paper at ICLR 2027

Localization Does Not Imply Steerability: A Mechanistic Study of Sign Errors in LLM Mathematics

Abstract

When large language models do multi-step mathematics, a characteristic failure recurs: the model writes one wrong sign token mid-computation — a cofactor sign, the minus of a recursive integration-by-parts step — then executes the rest of the arithmetic correctly, and the answer comes out simply wrong. We ask where inside the network that sign is written. Across six open models spanning 14B to 675B parameters, a per-layer readout shows the wrong sign is committed late — by MLP layers in the last quarter of the network (70–100% per model and task), a pattern already present, where tested (Qwen), in the base model before instruction tuning. The mechanism varies: some models carry fixed “minus-stamp” layers that push a minus regardless of the right answer, others carry amplifiers of whichever sign was already chosen, and Qwen adds a downstream rescue layer that opposes the stamp. The circuit is causal: subtracting a single sign direction from the residual stream flips 13 of 14 held-out minus-written integration-by-parts errors on Llama- 3.3-70B, with flips replicating on Qwen, Gemma, and Mistral-Small (24–70% per cell), and with representation-level specificity — the same dose breaks correctlyminus answers while sparing correct-plus answers where the direction is surgical (0/72 on Llama; Mistral, the outlier, breaks every cohort), and matched randomdirection and control-layer interventions are null. Core claims were preregistered with pass criteria frozen in git before held-out data was seen: 23 of 29 checks passed, all misses reported. But causality is not uniform: Phi-4 implements an equally well-localized circuit at the same late depth, yet the identical intervention flips almost nothing—the sign is written redundantly across ∼15 layers with margins several times larger (median 17.8 logits). Localization, we conclude, does not imply steerability: two models can share a circuit location yet differ enormously in how overturnable it is.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.