acceptodds
Under review as a conference paper at ICLR 2027

Spot the Difference: Detecting Hallucinations through Context-Head Intervention and Sparse Autoencoder Feature Contrasts

Abstract

RAG aims to enhance factuality in LLMs, yet these models frequently generate unfaithful outputs that contradict or exceed the provided context. While the internal representations of the LLM often contain features sufficient to flag hallucinations, the additive nature of the residual stream allows ingrained parametric knowledge to progressively outweigh contextual information during inference. In this work, we investigate this internal conflict, termed the Knowledge-Context Gap (KCG), in the sparse autoencoder (SAE) feature space. We show that jointly modeling paired SAE features provides complementary signals for hallucination detection, though constructing the paired representations via prompt-level manipulation incurs computational overhead and representational drift. To address these challenges, we propose RAGap, a mechanistic framework that characterizes KCG by isolating the LLM's parametric knowledge directly within the residual stream. By identifying and suppressing specialized context heads through differentiable masking, RAGap enables an efficient characterization of the latent divergence between parametric knowledge and contextual information. Leveraging SAEs for interpretable feature extraction, our method consistently outperforms state-of-the-art hallucination detectors across multiple benchmarks. Our findings provide a novel mechanistic perspective on RAG faithfulness and offer a practical solution for reliable, inference-time hallucination detection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.