Demystifying and Resolving the Pathologies of Linguistic Uncertainty in Large Language Models
Abstract
Across diverse languages, speakers naturally convey uncertainty through epistemic hedges before asserting claims; however, uncertainty quantification (UQ) in large language models (LLMs) relies almost exclusively on post-hoc numerical scores. Can LLMs learn human-like, verbalized uncertainty? To investigate this, we systematically analyze model behavior along two primary axes: confidence placement (pre- vs. post-assertion) and modality (numeric scores vs. linguistic hedges). Our analysis uncovers two critical pathologies: (i) pre-assertion confidence acts as an ungrounded prior that is statistically decoupled from the downstream assertion; and (ii) linguistic hedging is acutely brittle under prompt paraphrasing. Consequently, naive epistemic hedging severely degrades model reliability. To remedy these deficiencies, we propose Stratified, a framework that decomposes generation into distinct epistemic and linguistic tiers by first committing to a statement and its numeric confidence before verbalizing the hedged assertion. Furthermore, to elicit fine-grained hedging across long-form responses, we construct WikiEU, a benchmark spanning diverse, factually verifiable queries, alongside an optimization recipe combining confidence-balanced supervised fine-tuning and reinforcement learning with a gated reward mechanism. Across three benchmarks and two models, our approach substantially improves verbalized uncertainty expression without sacrificing factual fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.