acceptodds
Under review as a conference paper at ICLR 2027

From Detection to Mitigation: Neuron Attribution for Contextual Hallucination in RAG

Abstract

Retrieval-augmented language models may ignore externally provided evidence when it conflicts with parametric knowledge, leading to contextual hallucination. We introduce NADM (Neuron Attribution for Detection and Mitigation), a unified framework that uses response-conditioned neuron attribution as a shared internal signal for contextual hallucination detection and mitigation. NADM represents a generated response by computing gradient-based attribution from internal neurons to the logits of its realized tokens, and trains a simple linear classifier on the resulting attribution patterns to distinguish context-faithful from contextually hallucinated generation. Across multiple datasets and model scales, the NADM detector achieves strong discrimination between context-faithful and contextually hallucinated generation. We further examine the causal role of the key neurons identified by the detector and find that, although they participate in context–parametric competition, transferring their activations alone reproduces only a small fraction of the model's context-versus-parametric preference difference. This motivates using attribution as a readout rather than directly manipulating the selected neurons. At inference time, NADM samples alternative responses, scores them with the same frozen detector, and selects alternatives with stronger evidence of contextual faithfulness. This yields an inference-time intervention that requires no update to the base-model parameters and does not assume that direct manipulation of the selected neurons is sufficient to reproduce the desired source preference. Together, NADM connects detection and mitigation through the same response-conditioned attribution signal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.