acceptodds
Under review as a conference paper at ICLR 2027

Hallucinations Exhibit Shallower Entropy Valleys: Training-Free Hallucination Detection in LLMs

Abstract

Large language models often give answers that sound confident and coherent even when the answers are unsupported or false. This gap between fluent language and reliable content makes it difficult to know when a generated response should be trusted. Existing hallucination detectors usually inspect the completed text, compare it with external evidence, or generate multiple answers to measure disagreement. These approaches can be useful, but they also raise a sharper question: can hallucinations be detected from the model's own internal computation without additional models, retrieval, or repeated decoding? We study this question through large activation events: the small subset of neuron activations whose absolute values are unusually large at a given layer. Rather than treating these events only as a sparsity pattern, we track how often each coordinate emits them across a response and summarize their layer-wise profile with binary entropy. We smooth this profile, retain non-oscillatory local minima, and use the descent prominence of the dominant entropy valley as a training-free hallucination-risk signal without imposing a fixed layer partition. We observe that these events preserve much of the computation needed for reasoning under causal masking, while controls at the same sparsity do not. Empirically, on paired TruthfulQA answers, correct responses exhibit a clearer decrease-and-recovery entropy pattern across layers, and the resulting event-profile score achieves AUROC values of 0.76 for LLaMA-3-8B and 0.79 for LLaMA-3-70B. Valley-stage interventions localize the strongest causal sensitivity to the low-entropy core, while matched-sparsity threshold comparisons support the robustness of the event definition. These results suggest that internal event profiles across layers can serve as an interpretable diagnostic from a single forward pass for identifying responses that warrant further verification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.