Knowing What to Reconsider: Claim-Level Metacognition for Language Models
Abstract
Humans know that a claim can be false, outdated or unsupported, and they reconsider the ones that matter. Language models lack a good rule for which claims to re-check. On synthetic requests with a false premise among six claims, four open-weight models from 12B to 70B answer only 9–28% correctly. Yet two of their own signals usually flag that premise, a probe on hidden states and the model's probability that each claim is true. Naming that premise in a short note gets it re-checked; naming another claim does not. We propose to re-check a claim when its chance of being wrong, times how much the answer depends on it, times the odds a check catches it, exceeds the cost. We estimate that chance with either signal and hold the other terms fixed. If the most suspect claim passes a threshold, the note names it. With the false premise hidden among 24 claims, this policy lets the models answer 58–87% correctly, against 28–50% when asked to verify every claim the answer uses, while adding under half as many generated tokens. Catching a false claim takes not universal scepticism, but knowing what to reconsider.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.