How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
Abstract
Large language models can detect their own errors and sometimes correct them without external feedback, but the underlying mechanisms remain unknown. We investigate this through the lens of second-order models of confidence from decision neuroscience. In a first-order system, confidence derives from the generation signal itself and is therefore maximal for the chosen response, precluding error detection. Second-order models posit a partially independent evaluative signal that can disagree with the committed response, providing the basis for error detection. Kumaran et al. (2026) showed that LLMs cache a confidence representation at a token immediately following the answer (i.e. post-answer newline: PANL)—that causally drives verbal confidence and dissociates from log-probabilities. Here we test whether this signal supports error detection—the core prediction of the second-order framework—and a hypothesized extension to self-correction. Using a verify-then-correct paradigm, we show that: (i) verbal confidence predicts error detection far beyond token log-probabilities, ruling out a first-order account; (ii) PANL activations predict error detection beyond verbal confidence itself; and (iii) PANL predicts which answers the model will revise, and critically which errors the model can correct—beyond all behavioural signals. Causal interventions establish that PANL signals rescue error detection when answer information is corrupted, and that ablating answer-adjacent activations including PANL disrupts self-correction. All findings replicate across models (Gemma 3 27B and Qwen 2.5 7B) and datasets (TriviaQA, PopQA, and MNLI). These results reveal that LLMs naturally implement a second-order confidence architecture whose internal evaluative signal encodes not only whether an answer is likely wrong but whether the model has the knowledge to fix it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.