acceptodds
Under review as a conference paper at ICLR 2027

GLIB-SE: Low-Budget Semantic Uncertainty Quantification via Residual Semantic Mass Modeling

Abstract

Reliable deployment of large language models (LLMs) requires uncertainty estimates that remain informative under tight inference budgets. Existing semantic uncertainty estimators often rely on a limited set of sampled responses to characterize the semantic distribution, which can understate uncertainty when semantically plausible but unobserved alternatives still carry non-negligible probability mass. We propose GLIB-SE, which augments observed semantic classes with a Ghost Cluster to represent residual semantic mass. Semantic counts and energy derived from generated-token scores parameterize an augmented Dirichlet proposal, while response likelihoods constrain Monte Carlo entropy scoring. By incorporating model-internal information into the residual prior, GLIB-SE can distinguish response sets with identical semantic count distributions. We evaluate GLIB-SE using ten responses per input generated by nine language models across six datasets. GLIB-SE achieves higher mean AUROC than likelihood-weighted and count-based semantic entropy in all 18 dataset–model combinations and the highest mean among seven methods in 16 combinations. Support-stratified analysis shows that most of the gain over count-based entropy arises when all sampled answers fall into one observed semantic class, where counts alone provide no ranking information. These findings support combining residual semantic support with model-internal signals to improve error detection under limited sampling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.