Tokens, Numbers, or Words? Benchmarking How Large Language Models Express Uncertainty
Abstract
As large language models (LLMs) are increasingly deployed in high-stakes applications, robust uncertainty estimation is essential for ensuring the safe and trustworthy deployment of LLMs. We present a large-scale study of uncertainty estimation in LLMs, evaluating 80 models spanning open- and closed-source families, dense and Mixture-of-Experts (MoE) architectures, reasoning and non-reasoning modes, quantization variants and parameter scales from 0.6B to 671B. Focusing on three representative single-pass methods that require no access to hidden states, including token probability-based uncertainty (TPU), numerical verbal uncertainty (NVU), and linguistic verbal uncertainty (LVU), we systematically evaluate uncertainty calibration and selective classification using the challenging MMLU-Pro benchmark, which covers both reasoning-intensive and knowledge-based tasks. Our results show that, on MMLU-Pro, LVU outperforms TPU and NVU on average and for most models, offering stronger calibration and discrimination. We also find that high accuracy does not imply reliable uncertainty, and that model scale, post-training and reasoning ability influence estimation performance, whereas the effects of quantization are small and method-dependent. Notably, LLMs exhibit better uncertainty estimates on reasoning tasks than on knowledge-heavy ones, and calibration and error ranking are only partially aligned. These findings highlight the need for multi-perspective evaluation and position LVU as a promising, though not universally superior, tool for improving the reliability of LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.