acceptodds
Under review as a conference paper at ICLR 2027

Calibrating Verbalized Confidence with Rubrics

Abstract

Large language models (LLMs) can assess the correctness of their own responses through verbalized confidence, helping users decide whether to adopt the response or seek further review. However, models often exhibit overconfidence when their responses are incorrect, limiting the practical value of these estimates. Existing research addresses this problem by refining how confidence judgments are elicited and aggregated, yet whether expressed confidence is reliable under more stringent and explicit verification criteria remains underexplored. In this paper, we systematically analyze models' verbalized confidence through progressively stricter rubric-based checks and propose two confidence calibration methods. The proposed methods integrate verification signals from these checks with the model's holistic self-assessment, thereby mitigating overconfidence. Experiments on eight benchmarks spanning diverse reasoning tasks and domain knowledge, across three model families and three model scales ranging from B to B, demonstrate improved calibration while preserving discrimination between correct and incorrect responses. Further analysis shows that rubric-based verification provides useful signals beyond direct confidence, and that the gains cannot be attributed merely to more conservative confidence estimates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.