What Matters, Not When: Deciphering Prompt Reuse with Distributional Profiling for Hallucination Detection in LLMs
Abstract
Although large language models (LLMs) have demonstrated remarkable reasoning capabilities across diverse domains, they remain prone to hallucinations that undermine their reliability. To detect hallucinations, existing approaches generally leverage hidden states or attention matrices produced during generation as discriminative signals. Despite their effectiveness, these methods often implicitly assume that hallucination evidence is primarily encoded in generated tokens and their sequential organization. However, our empirical analysis reveals that hallucination risk can already be predicted before decoding, while detection remains largely insensitive to explicit temporal ordering. Motivated by these observations, we propose a novel framework, termed Distributional Profiling of Prompt Reuse (DOOR), for hallucination detection in LLMs. The core idea of DOOR is to characterize prompt reuse through the attention between prompt and generated tokens along two complementary dimensions: reuse intensity and reuse dispersion. More specifically, we first construct a distributional profile of prompt reuse from the attention received by prompt tokens during generation. Subsequently, we leverage mean reception and coefficient of variation to quantify reuse intensity and reuse dispersion, respectively, abstracting away from the temporal order of generation. To accommodate prompts of varying lengths, DOOR further introduces a rank-based profiling module that yields fixed-dimensional features, which are subsequently fed into a logistic regression probe for hallucination detection. Extensive experiments across diverse model backbones and benchmarks demonstrate the effectiveness and generalizability of DOOR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.