Right Answer, Wrong Reasoning: Localizing Chain-of-Thought Restoration Errors in the Residual Stream
Abstract
Chain-of-thought traces are increasingly treated as explanations, yet that trust depends on whether the stated reasoning faithfully reflects how the model reached its answer. We study a failure that escapes surface-level checks: the model reaches the correct answer while producing a verifiably wrong intermediate step that is fluent and never corrected. We call this silent restoration. Answer-level verification misses it because the answer is correct. Across twelve model configurations and four domains, silent restoration appears universally and reaches 32% of answer-correct MedQA traces for Qwen2.5-7B. Across nine model configurations, linear probes identify wrong steps at 0.73–0.91 AUROC, consistently at interior layers. We also find that this signal is already present before the step is generated. Across layers, wrong steps exhibit a uniform upward shift in probe score rather than a change in profile shape. Exploiting this geometry with a probe-plus-distance detector yields 0.792 out-of-domain AUROC, outperforming all baselines and the strongest token-uncertainty signal at 0.655. The signal is also causal: steering along the probe direction induces verifiable errors in 46% of correct steps, and can also repair naturally wrong steps. The same intervention distinguishes load-bearing steps, on which the answer depends, from performative ones, on which it does not, suggesting that faithfulness is fundamentally a per-step rather than a per-trace property.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.