When Measures Become Targets: Construct Validity Failures in Uncertainty-Based Hallucination Detection
Abstract
Uncertainty-based hallucination metrics use token confidence, sampled semantic agreement, and representation geometry as proxies for factual reliability. We investigate when these signals diverge from factuality or semantic consistency under controlled perturbations. We develop metric-specific diagnostic tests for First Output Token Entropy, Semantic Entropy, and EigenScore, and evaluate them across four question-answering datasets and four model configurations from three model families. Targeted prompt edits lower first-token entropy while incorrectness is preserved; factual corruption remains below the NLI-based Semantic Entropy threshold; and meaning-preserving paraphrases can increase EigenScore enough to trigger hallucination flags. A human audit supports these discrepancies in the assessed examples and shows that low NLI-estimated entropy can coexist with reduced perceived semantic consistency. We also find related discrepancies under fixed generation, prompting, and rewriting conditions without metric-directed optimization. We therefore identify limitations of the evaluated operationalizations and motivate complementing predictive benchmarks with controlled tests of factuality and semantic preservation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.