Dispersion Is Not Doubt: Rethinking LLM Confidence Estimation When Multiple Answers Are Valid
Abstract
Confidence estimators often treat concentrated support as evidence of correctness, relying—explicitly or implicitly—on the assumption that concentrated support implies a correct answer. When several answers are valid, however, this assumption fails: support can spread across correct alternatives without indicating error. We call this mismatch the concentration–correctness gap, and propose two layered conjectures. (a) The inversion conjecture: as the number of valid answers grows, confidence falls even as the model answers better. (b) The dispersion-is-not-doubt conjecture: the decline stems from support being re-divided among valid answers, not from greater uncertainty about correctness. To test them, we introduce MACE (Multi-Answer Confidence Estimation), open-ended questions with controlled answer cardinalities () and exhaustive, expert-audited reference sets. Across 9 models and 15 training-free estimators, both conjectures are confirmed: a multi-answer confidence inversion is widespread; all 15 estimators lose AUROC under pooled cardinalities; and valid mass stays stable on correctly answered questions while support is unevenly re-divided—entropy decomposition, elicited confidence, and cross-model analysis () jointly attribute the decline to valid-answer diversity. These findings motivate CARE (CArdinality-Rescaled Evidence), which rescales sampled answer frequency by question-level cardinality to approximate the total support on valid answers. Using verified counts, CARE raises AUROC from to on the broadest mixture, recovering approximately of the loss and demonstrating the value of estimating the number of valid answers for more reliable confidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.