Evidential Semantic Uncertainty for Autoregressive LLM Generation
Abstract
Uncertainty quantification (UQ) is essential for improving the reliability of autoregressive generation in large language models (LLMs). In open-ended generation, uncertainty can arise at lexical, syntactic, and semantic levels, among which semantic uncertainty is most indicative of answer correctness. A widely used approach, semantic entropy, estimates uncertainty by sampling multiple answers, clustering them into semantically consistent groups, and computing the entropy over group probabilities. However, this formulation assumes disjoint answer groups and fails to capture partial semantic overlap. Recent methods address this limitation by leveraging spectral clustering and kernel-based representations to model dispersion in semantic space. Despite these advances, an important failure mode remains unaddressed: cases where the model produces consistently incorrect yet semantically coherent answers. In this work, we observe that such incorrect generations are often associated with low evidential support, providing a complementary signal for uncertainty estimation. Building on this insight, we propose an evidential-theoretic framework that models answer groups using multinomial opinion and hyper-opinion. The proposed framework enables a principled decomposition of semantic uncertainty into: (i) lack of supporting evidence and (ii) conflict among competing answer groups. Our formulation explicitly captures inter-group semantic overlap and yields more expressive uncertainty measures aligned with the underlying semantic structure. Experiments across multiple question-answering benchmarks and LLMs demonstrate that our method consistently improves uncertainty estimation and enhances the detection of incorrect answers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.