Beyond Binary Correctness: Confidence Estimation for Continuously Graded Responses in Large Language Models
Abstract
Confidence estimation predicts the correctness of generated outputs of LLMs. Most existing approaches typically use an arbitrary threshold over a continuous quality metric, such as syntactic or semantic similarity against ground truth, to determine whether an LLM response is correct. In this paper, we propose modeling confidence as a probability distribution over continuous correctness scores, thereby remaining agnostic to threshold choice, avoiding information loss from binarization, and enabling rich uncertainty estimates such as confidence variance and credible intervals for robust decision making. We propose Beta regression and piecewise uniform regression approaches as distributional methods, showing theoretically that Beta regression achieves superior sample efficiency and faster convergence for smooth distributions with at most one interior mode. We also conduct an empirical study on several datasets across diverse tasks to analyze and compare methods with each other, as well as traditional point-based confidence estimation baselines, demonstrating their relative suitability and advantages.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.