acceptodds
Under review as a conference paper at ICLR 2027

The Circuit Basis of Metacognitive Failure in Language Models: Conflict and Hallucination

Abstract

Language models often answer with high confidence, as measured by the entropy of the output distribution, even when they are wrong. We study two situations that lead to such errors, *knowledge conflict*, where the context contradicts a fact stored in the weights, and *missing knowledge*, where the queried fact was never learned and the model *hallucinates*. To control exactly what the model has learned, we fine-tune it on synthetic facts that map arbitrary entity identifiers to random codes, installed through low-rank adapters (LoRA) on the multi-layer perceptron (MLP), query-key (QK), or value-output (VO) components. Our analysis suggests that MLP is the main substrate for storing facts in *basins*, regions of representation space into which a fact's representations converge, while QK steers generation toward a basin or the context, and VO writes the retrieved content into the residual stream. *Under knowledge conflict*, the arbitration between the stored fact and the context fails in a way that depends on how the fact is stored. A fact learned from a single phrasing yields to the context, but errors accumulate in the later digits, while a fact learned from many phrasings overrides the context, and the answer drifts into one that matches neither source. *Under missing knowledge*, the hidden state reaches no basin, yet every adapter placement produces confident hallucinations, most of all through VO. In both situations, the output entropy reflects how sharply the model commits rather than whether the right fact was retrieved, even though this information remains present in its representations. Across 12 pretrained models, the overall hallucination rate falls with scale, but a growing fraction of the remaining hallucinations are delivered with high confidence. Our mechanistic insights thus pinpoint a systemic metacognitive gap that does not close with scale in the models we study.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.