acceptodds
Under review as a conference paper at ICLR 2027

Don’t Be Choosy: Scoring over Choosing for Verbalized LLM Confidence

Abstract

Accurate uncertainty estimation is critical for the reliable deployment of large language models (LLMs). A simple and widely used approach is to ask a model to verbalize its confidence on each item, referred to as AbsScore henceforth. However, AbsScores can be poorly discriminative, exhibiting overconfidence or collapsing to a small set of values. Recent work proposed pairwise elicitation as an alternative, asking the model which of two items it is more confident about. While effective, this formulation forces a preference even when confidences are similar and discards information about the magnitude of the confidence difference. To overcome these limitations, we introduce PairScore, a simple alternative: present two items together, but ask the model to score both rather than choose between them. This preserves graded confidence while retaining the benefits of pairwise elicitation, and the resulting scores can be aggregated directly without fitting a separate ranking model. PairScore improves confidence ranking over AbsScore and PairChoice in most claim-verification settings, while gains in best-of- response selection depend on the task. PairScore also improves calibration and reduces sensitivity to presentation order and comparison partners. We extend both formats beyond pairs: -score jointly scores items and -rank returns a full ordering of them. Scoring improves as more items share a prompt, whereas ranking peaks at small and then declines; -score matches or exceeds -rank at every tested.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.