acceptodds
Under review as a conference paper at ICLR 2027

Epistemic Uncertainty in Large Language Models: Measurement, Evaluation, Mechanisms, and Applications

Abstract

Token-level uncertainty in large language models (LLMs) arises from both knowledge gaps (epistemic uncertainty, EU) and the inherent flexibility of language (aleatoric uncertainty, AU). Accurate token-level EU estimation is fundamental to understanding model reliability at the basic decision unit of generation. However, the lack of a unified and direct testbed leaves estimator effectiveness, EU signal mechanisms, and downstream applicability insufficiently understood. To advance the understanding of uncertainty in LLMs, we make four contributions: 1. We introduce EU and AU proxy labels based on reasonable continuation sets and develop an automated pipeline to construct UncertainBench, comprising 7,216 labeled continuation positions across 15 domains. 2. We systematically evaluate EU across 67 LLMs and benchmark 35 uncertainty estimation methods, revealing that none combines accurate EU estimation with high EU purity. 3. We discover stronger and purer EU signals in the low-probability tail and propose a training-based hypothesis consistent with observed logit patterns. Emphasizing these signals yields a training-based estimator that outperforms baselines and generalizes across domains. 4. We apply EU estimators to five downstream tasks, helping distinguish reliable knowledge from lucky success, assess response quality, predict model capability, and identify critical branching positions. We also find that EU estimators offer limited value for predicting pre-trained model potential, whereas our proposed R\'enyi entropy-based method performs well by emphasizing tokens in head probability buckets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.