acceptodds
Under review as a conference paper at ICLR 2027

Dissociating Internal Error Detection from Stated Verdicts in Small Language Models

Abstract

Small instruction-tuned models asked whether an arithmetic derivation is correct answer yes almost regardless of the derivation. We show that the evidence of the error is present in the model's own next-token probabilities and, at 1.5B and above, in the state its verdict is read from, and that the verdict uses little of it. On 1,280 seeded derivations with one planted wrong step, the model's own next-token surprisal separates corrupted from clean derivations with AUC 0.872 (Qwen2.5-0.5B-Instruct) and 0.893 (Qwen2.5-1.5B-Instruct) and finds the wrong step 75.2% and 79.2% of the time, while the Yes/No verdict on the same inputs reaches AUC 0.507 and 0.544. We split verification into three stages measured on the same problems: noticing the error when it is read, carrying the evidence forward, and reporting it in the answer. The 0.5B model loses the evidence within one step. In the 1.5B model, a linear probe on the final normalized state at the answer position, the vector the unembedding reads, detects corruption with AUC 0.803 regardless of the error's distance from the question, and the corruption direction is nearly orthogonal to the No-minus-Yes unembedding (cosine 0.069). On a subset of short derivations, removing that component lowers the verdict's AUC from 0.663 to 0.558 and amplifying it tenfold raises it to 0.787, while a random direction moves it by at most 0.034. The pattern persists at 3B and 7B, where surprisal reaches AUC 0.920 and 0.931 and the verdict 0.539 and 0.605; carry and the final-state probe vary with size and weight precision rather than growing with size. On ProcessBench, the digit surprisal of Qwen2.5-Math-1.5B-Instruct localizes the first error with an average F1 of 42.9, with no training and one threshold chosen on GSM8K, level with a 7B process reward model trained on step labels, while the same model asked for the index scores 0.7.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.