Detecting Contextual Hallucinations by Pooling and Probing Attention-Head Outputs
Abstract
Retrieval-augmented language models can produce claims that are unsupported by their supplied context, motivating methods for detecting contextual hallucinations. Prior work has shown that truthfulness and hallucination can be predicted from internal model representations, suggesting that these failures leave detectable signatures within the model. Building on linear probing methods for truthfulness, we pool each attention head's current-value contributions across generated positions into a response-level representation and train a linear probe to learn a hallucination-discriminating axis. Across multiple RAG hallucination benchmarks and model families, HeadAxis consistently matches or exceeds strong white-box hallucination detectors. Analysis shows that the axes capture a hallucination signal that generalizes across datasets and can be decomposed across response tokens without additional supervision, while ablations show that the representation is robust across alternative design choices.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.