SHEAR: Hidden-State Editing for Mid-Thought Reasoning Correction
Abstract
Large reasoning models (LRMs) generate extended chains of thought, but do they internally know when they have already found the correct answer? We investigate this question by probing hidden states during reasoning. A lightweight probe on hidden states identifies per-step reasoning correctness with high accuracy and shows a sharp jump at the moment of answer stabilization, providing strong affirmative evidence. This information is substantially more accessible through hidden states than through token-level signals: in controlled equal-dimension comparisons using the same classifier, hidden-state probes markedly outperform token-level probes. We propose SHEAR (Semantic Hidden-layer Editing with Anchored Reasoning), which directly exploits this hidden knowledge through a semantic confidence probe and a stability-anchored editing direction that jointly determine when and how to intervene. SHEAR achieves the highest or tied-highest results on six mathematical reasoning benchmarks across three LRM scales, with statistically significant improvements against the strongest baseline in most configurations. Reversing the editing direction causes catastrophic collapse, indicating that it operates on genuine reasoning structure. Hidden states are therefore not merely a readout of reasoning quality but a control surface for correcting it mid-thought. Code is available at https://anonymous.4open.science/r/SHEAR-C9A4/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.