acceptodds
Under review as a conference paper at ICLR 2027

Verbalized Confidence Calibration for Retrieval-Augmented Large Language Models via Multi-Source Evidence Fusion

Abstract

Large language models are increasingly deployed in retrieval-augmented generation (RAG) settings, yet they frequently verbalize high confidence when their answers are incorrect. Existing methods for calibrating verbalized confidence typically rely on token-level generation probabilities or heuristic confidence scores, and use the correctness of the predicted answer as supervision, without explicitly modeling the agreement and conflict across the model's parametric knowledge and the retrieved documents. We propose a two-stage framework for verbalized confidence calibration in retrieval-augmented multiple-choice question answering. In the first stage, information from the model's parametric knowledge and from each retrieved document is encoded into source-specific evidence vectors, each of which is used to parameterize a Dirichlet distribution over answer options. Using self-attention, a fusion module captures agreement and conflict among these sources and produces a fused Dirichlet distribution. The expected value of the probability assigned to the predicted answer under this distribution serves as a scalar confidence target. In the second stage, a lightweight confidence estimator takes the question-conditioned answer-option logits as input and learns to predict this scalar target without requiring additional document retrieval during confidence estimation. The predicted scalar confidence is then converted into a Gaussian-smoothed soft target distribution over verbalized confidence levels, and KL divergence regularization is employed to align the model's verbalized confidence distribution with this soft target. Experiments on multiple-choice question-answering benchmarks show that our method improves confidence calibration on the in-distribution dataset and on most out-of-distribution datasets, and, overall, yields more reliable uncertainty estimates than baseline methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.