acceptodds
Under review as a conference paper at ICLR 2027

ACCOUNTING FOR THE UNSEEN: FRAGILITY-AWARE RECALIBRATION OF SEMANTIC UNCERTAINTY FOR HALLUCINATION DETECTION

Abstract

Despite their impressive ability to generate natural and coherent text, large language models (LLMs) can produce plausible but fabricated responses. The same prompt, when repeatedly applied, may yield different answers, undermining the confidence one may place in an individual generation. The common approach of detecting such confabulations by observing multiple generations to estimate the uncertainty over their semantic content must necessarily rely on a finite sample of all generations that the model may produce. Such estimates can be fragile when rare or singleton answers dominate observed semantic support. To address this, we propose Semantic Uncertainty Recalibration accounting for Fragility (SURF), a framework that corrects semantic uncertainty estimates using the probability of observed singletons as a proxy for unobserved semantic support. SURF incorporates this fragility proxy into an entropy-based optimization that recalibrates token-sequence probabilities according to their sensitivity, yielding uncertainty estimates that are more robust against finite observations. We also provide an estimate of the number of semantically distinct answers as an actionable signal for when additional observations or human intervention may be warranted. When evaluated across 112 different settings (involving four datasets, seven model variants, and different generation lengths and model quantization levels), SURF shows consistent improvement in AUROC and headroom gain (HRG), and maintains stable AURAC performance, which demonstrates its reliability even under resource constraints.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.