acceptodds
Under review as a conference paper at ICLR 2027

Reliability and Informativeness of Verbalized Confidence in Language Models

Abstract

Language models are increasingly asked to state how confident they are, and the numbers they give back are reported as calibration evidence. Calibration presupposes a measurement, however: a quantity that carries information about correctness and that stays put under changes to how it is requested. We test both properties across open-weight instruction-tuned models spanning more than an order of magnitude in scale, on five English benchmarks, under five prompt formulations and seven ways of asking for a confidence value. Verbalized confidence separates correct answers from incorrect ones only slightly better than chance, with a mean AUROC of 0.536 against 0.682 for the first-token probability of the same responses. It is also unstable. Holding the model, the question and the model's own answer fixed and changing only the wording of the request, the two formulations used in prior work barely agree, and a model asked how likely its answer is to be wrong does not give a value consistent with what it reports when asked how likely that answer is to be right. Often no value is recorded at all, and much of that missingness belongs to the extraction rule rather than to the model. None of this resolves with scale: within one family at fixed training recipe, accuracy and first-token informativeness both rise steeply across a sixty-six-fold parameter range while the verbalized report stays near chance. A model trained to verbalize confidence behaves differently under the same procedure, which indicates that our measurement registers a trained signal where one is present. The practical consequences follow. Which model looks best calibrated changes with the wording of the request in roughly a third of our comparisons, and abstention built on this signal recovers about a seventh of the accuracy available from token probability. We recommend that checks of informativeness and reliability accompany any reported verbal calibration number.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.