acceptodds
Under review as a conference paper at ICLR 2027

Heterogeneous Uncertainty Along Denoising Traces Signals Hallucinations in Diffusion Language Models

Abstract

Given the rich token-confidence information they contain across positions and time steps, the denoising traces of diffusion language models offer a natural foundation for uncertainty quantification and hallucination detection. Yet, existing detection methods either derive from autoregressive models—failing to capture the full extent of this information—or process it using heavyweight neural architectures. In this paper, we show that hallucinations are better predicted by the heterogeneity of uncertainty along this trace than by its average level. In particular, large variations in token distribution entropy over denoising steps provide a strong discriminative signal. A Markovian model of token-wise convergence links this pattern to slower and more heterogeneous convergence in hallucinated answers. Building on these findings, we introduce Drifty: a lightweight detector combining 37 closed-form trace features with regularized logistic regression. Across three diffusion language models and three closed-book question-answering benchmarks, Drifty outperforms all evaluated baselines in every model–dataset pair, improving over Semantic Entropy by up to 14.5 ROC-AUC points and 16 PR-AUC points. Moreover, Drifty uses only the trace of a single deterministic generation, with feature extraction and classification requiring roughly half a second of CPU time per 100 samples.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.