Language Models Do Not Re-verify Computed Values in Their Own Reasoning
Abstract
Language models exhibit increasingly remarkable capabilities on reasoning tasks, but their robustness to errors within their own reasoning remains uncertain. In this paper, we investigate how models respond to small perturbations of their reason ing traces. In particular, we inject a single numerical error into model-generated execution traces for short mathematical programs across 14 language models from 5 developers, and study whether the models detect and correct the error or continue reasoning from it. We find that error detection is largely determined by how the value is used in the model. Specifically, when the planted value explicitly contradicts a value written earlier in the trace, detection and correction raises up to 96 − 99% for stronger models. In contrast, when detecting the error requires re-executing even a single prior operation, every model capable of solving the task generally fails at detecting and correcting the value, with a success rate of 0−12% in the best case. Importantly, models in those setups are able to provide the value when asked directly. This failure persists even in cases where verification instructions are given, and the necessary information and computational capability are available, and is also not repaired by the activation-steering interventions we test. Our results expose a fragility in downstream self-correction even in highly capable models, and challenge the assumption that an incorrect intermediate computation will have a meaningful chance of being independently rechecked later in the rea- soning process
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.