acceptodds
Under review as a conference paper at ICLR 2027

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

Abstract

Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Yet existing methods estimate confidence in a new answer from the current inference alone, leaving the model's record of past successes and failures unused. We propose XConf (eXperiential Confidence): estimating confidence from the model's accumulated experience. XConf maintains a memory of graded past episodes, each recording the task, the stated confidence, the outcome, and a lesson drawn from the feedback. Given a new task, XConf first Recalls similar episodes to estimate the model's historical success rate and then Reflects on the retrieved outcomes and lessons to identify relevant failure patterns and revise its confidence in the current answer. Our estimator is format-general, requiring no logit access or weight updates, and costs only one answer generation. Across nine benchmarks spanning reasoning, coding, multimodal QA, and interactive agents, and four models from three families, XConf outperforms or matches ten-sample self-consistency in discrimination (AUROC) on 23 of 24 comparisons, with much lower calibration error (ECE), at a tenth of the generation cost. Used for selective prediction, abstaining on the 10% least-confident episodes raises the delivered success rate by up to 8.7 points on agent tasks. We therefore see experiential confidence estimation as a new paradigm for future general-purpose confidence estimation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.