HDIR: Training-free Hallucination Detection via Discriminative Internal Representations
Abstract
Large language models have achieved remarkable success in natural language processing, but they remain prone to hallucinations, generating plausible yet factually incorrect content. Effective hallucination detection is therefore crucial for assessing the factuality of generated content. Existing methods mainly rely on output signals or internal representations: the former are efficient and training-free but overlook richer semantic information encoded in model states, while the latter can better exploit such information but often require additional training. To address these limitations, we propose Hallucination Detection based on Discriminative Internal Representations (HDIR), a training-free framework that exploits internal representations for hallucination detection. Our analysis shows that the norm of attention-weighted internal representations provides a strong discriminative signal for hallucination detection. We define this measure as Semantic Intensity and further propose a Cohen's -based selection strategy to identify the most discriminative attention heads and dimensions, thereby strengthening the discriminative signal of Semantic Intensity. Extensive experiments across multiple benchmarks and LLMs demonstrate that HDIR achieves strong overall hallucination detection performance, generalizes well across datasets, and complements output-level uncertainty signals, with their combination yielding further performance gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.