Spot the Difference: Detecting Hallucinations through Context-Head Intervention and Sparse Autoencoder Feature Contrasts
Abstract
RAG aims to enhance factuality in LLMs, yet these models frequently generate unfaithful outputs that contradict or exceed the provided context. While the internal representations of the LLM often contain features sufficient to flag hallucinations, the additive nature of the residual stream allows ingrained parametric knowledge to progressively outweigh contextual information during inference. In this work, we investigate this internal conflict, termed the Knowledge-Context Gap (KCG), in the sparse autoencoder (SAE) feature space. We show that jointly modeling paired SAE features provides complementary signals for hallucination detection, though constructing the paired representations via prompt-level manipulation incurs computational overhead and representational drift. To address these challenges, we propose RAGap, a mechanistic framework that characterizes KCG by isolating the LLM's parametric knowledge directly within the residual stream. By identifying and suppressing specialized context heads through differentiable masking, RAGap enables an efficient characterization of the latent divergence between parametric knowledge and contextual information. Leveraging SAEs for interpretable feature extraction, our method consistently outperforms state-of-the-art hallucination detectors across multiple benchmarks. Our findings provide a novel mechanistic perspective on RAG faithfulness and offer a practical solution for reliable, inference-time hallucination detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.