Confidence Without Control: Revision in Language Models Is Confidence-Blind
Abstract
When a user pushes back on a correct answer, language models often give it up. Qwen2.5-7B answers of our held-out questions correctly, and after one sentence, "Actually, another source says the answer is X," its accuracy is . We show that this failure is confidence-blind revision: changing the model's stated confidence, with its answer held fixed, does not change whether it revises. The models know which answers are likely wrong, since stated confidence, P(True), and belief (the log-probability of the answer) predict their correctness with AUROC up to . Yet a perfect confidence note written into the model's own turn ( if right, if wrong) changes accuracy after the challenge by at most in three models from two families, although the models read the note and recall it on at least of items. A wrong-for-wrong test, which challenges a wrong answer with a different wrong answer, shows that Qwen2.5-7B and Llama-3.1-8B switch to the option they believe least as readily as to their second choice. Qwen2.5-14B, Qwen2.5-32B, and Qwen3-8B compare the two options more, and two frontier models abandon far fewer correct answers and often move to the true one. Training makes revision confidence-aware: trained models given a perfect note keep to of their answers correct under challenge, and models trained to follow a confidence threshold realize to of the best achievable gain (three seeds) for notes with AUROC of or more. Ignoring confidence costs accuracy in two ways. By default, models that are usually right abandon answers under an uninformative challenge, and a threshold rule on their own calibrated confidence, which keeps every answer, beats them by up to points. Item by item, a note with AUROC , acted on as well as possible, would leave Qwen2.5-7B at under challenge (it ends at ), but the models' self-report (AUROC to on these items) is too weak to add more than a point beyond keeping every answer. Calibration research should test whether models act on what they know.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.