acceptodds
Under review as a conference paper at ICLR 2027

Hidden Doubt, Missed Correction: A Causal Pathway in Chain-of-Thought Reasoning

Abstract

The errors made by large language models (LLMs) in chain-of-thought (CoT) reasoning are often decodable by linear probes from their activations. Yet the error signals in the representations do not always cause the model to self-correct. We study the gap between error awareness and error correction in LLMs. The Jacobian lens (J-lens) is an interpretability tool that maps activations to tokens the model is poised to verbalize. The mapping uses an averaged Jacobian from the activation to the final layer, which aggregates downstream paths through the model's computation graph across layers and subsequent token positions. Using the J-lens, we find that erroneous CoT steps exhibit stronger representations of doubt-related tokens, such as "But" and "However". Yet these tokens receive little probability mass in the next-token distribution. To verify whether the hidden doubt causally triggers self-correction, we steer activations along the corresponding J-lens directions, which makes the model voice the doubt and revise the reasoning path, improving final-answer accuracy. We also steer along the probe direction to amplify the error signals and find that this amplification can support self-correction by eliciting doubt. The two intervention approaches support hidden doubt as a possible mechanistic link between error awareness and self-correction. We further compare against non-doubt token interventions and direct text insertion, supporting that the effect is specific to eliciting doubt.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.