Token-FAR: Hallucination Detection from a Token Local Geometry Perspective
Abstract
Hallucination remains a major obstacle to the reliable deployment of Large Language Models (LLMs). However, existing hallucination detection methods often rely on repeated generation, semantic comparison, or trained detectors, causing extra inference cost or requiring supervision. In this paper, we investigate whether the local geometry of the internal token representations from a single candidate response can provide an effective hallucination signal without the need for training. To this end, we propose Token-FAR (Token-level Factual Affine Reconstruction), a token-level geometric hallucination detector based on local affine reconstruction with respect to a factual reference dictionary. To be specific, in Token-FAR, we utilize the hidden states of factual answers to construct a factual reference dictionary, and then for each candidate content token, we compute an affine ridge reconstruction residual with respect to a local factual neighborhood. We observe that the content tokens in a candidate answer that are poorly reconstructed by the local factual neighborhood will result in larger residuals, which provide a signal to indicate an occurrence of hallucination. By aggregating the token-level residuals, we define a score for detecting answer-level hallucination. Our proposed Token-FAR requires only a single candidate response at inference time and does not need to train a parametric detector using hallucination labels. We conduct extensive experiments and ablation studies on four factuality benchmark datasets across two LLMs. Experimental results demonstrate the effectiveness of our proposed approach, showing that local affine reconstruction with respect to a factual reference dictionary can provide an effective geometric signal for hallucination detection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.