Hallucination Detection in Diffusion Language Models via Sparse Token Informativeness
Abstract
Diffusion large language models (DLMs) have recently emerged as a promising alternative to autoregressive LLMs, with recent systems scaled to as many as 100B-parameter. Yet hallucination in DLMs remains underexplored: we still lack an interpretable understanding of where factual errors emerge during remasking, undermining their reliability in real-world applications. Existing DLM-specific detectors recover richer denoising evidence by training supervised models over selected sub-traces or structured token-interaction graphs. We identify an Early Factual Commitment phenomenon in DLMs: factual responses tend to become distinguishable before generation is complete, as decisive token commitments reduce uncertainty early in the denoising trajectory. Motivated by this observation, we separate each intermediate state into two views: the already committed stable context and the remaining masked positions that encode predictive uncertainty. For each token commitment, we measure its predictive information gain by tracking how it changes uncertainty over these two views. Extensive experiments show that sparse commit-level uncertainty signals provide effective yet interpretable hallucination detection while requiring no auxiliary verifier, repeated sampling, dense trajectory modeling, or learned detector architecture.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.