SpanProbe: An Adaptive Hallucination Detection Framework Based on Token Evidence Profiles
Abstract
Multimodal large language models (MLLMs) can generate fluent visual descriptions but remain prone to hallucinations, limiting their reliable deployment. Existing hallucination detection methods based on model-internal hallucination signals typically rely on predefined readout strategies. Because tokens within a span carry different information, fixed readout strategies struggle to adequately capture the relative importance of individual tokens, making it difficult to reliably assess the hallucination risk of a complete semantic span. We find that hallucination signals tend to concentrate on a small number of tokens within each span and that the token carrying the strongest signal varies across samples. Therefore, accurately characterizing the relative importance of individual tokens for hallucination discrimination under such a non-uniform distribution is essential for reliable span-level hallucination detection. To address this challenge, we propose SpanProbe, an adaptive hallucination detection framework based on Token Evidence Profiles (TEPs). SpanProbe extracts visual-attention mass, full-attention entropy, and attention-head output norms from every attention head in every decoder layer to characterize each token's visual dependence, attention dispersion, and activation strength. It then dynamically generates sample-dependent aggregation weights from the token TEPs and uses these weights to construct a span-level representation for hallucination detection. Furthermore, SpanProbe employs a risk-gated local rollback-and-regeneration strategy. Using the risk scores produced by the pooling module, it identifies high-risk semantic spans and selectively revises high-risk content while preserving previously accepted low-risk content, thereby mitigating hallucinations during generation. Extensive experiments across multiple benchmarks demonstrate that SpanProbe generally outperforms existing baselines. On M-HalDetect, which contains multi-token semantic spans, SpanProbe achieves an average AUPRC of 77.43%, outperforming the strongest baseline by 8.05 percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.