acceptodds
Under review as a conference paper at ICLR 2027

What Does Semantic Uncertainty Reveal and What Does It Miss? A Large-Scale Study of Semantic Uncertainty in LLM Natural Language Generation

Abstract

Semantic uncertainty quantification (SUQ) has become a central paradigm for detecting hallucinations in free-form language generation, as it measures uncertainty over meanings rather than surface forms. However, existing studies typically evaluate SUQ under limited tasks, models, and generation settings, making it difficult to understand when semantic uncertainty is reliable and where it fundamentally fails. We investigate both questions through SURE, a large-scale empirical study treating SUQ as both an object of evaluation and a lens on LLM behavior. We compare 19 measures and variants from six studies across four tasks, 19 English-output datasets, and 13 LLMs, complemented by 12-model translation experiments across 14 non-English-target directions. First, we analyze Semantic Uncertainty itself: method rankings vary with task and simple consistency measures remain competitive in several evaluated settings. We identify SURE-OC, diagnostic subsets of low-quality generations assigned confidence, to characterize where semantic agreement and quality diverge. Second, we use semantic uncertainty to analyze LLMs: higher-performing models tend to exhibit lower uncertainty, including on low-quality outputs, while QA temperature sweeps show that better uncertainty separability can accompany worse answer accuracy. Together, the two perspectives distinguish the effectiveness of an uncertainty estimator from the behavior of the generator, revealing why lower semantic uncertainty alone does not establish greater reliability. .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.