Suppression Is Not Correction: Causal Tracing of RAG Corpus Poisoning
Abstract
Retrieval-augmented generation (RAG) can be poisoned by malicious corpus passages that steer models toward answers chosen by attackers. Existing studies show that these attacks are effective, but do not explain how poisoned evidence changes the model internally or why removing its influence may fail to recover the correct answer. We present a causal tracing framework that follows retrieved evidence from document positions through the residual stream to the final answer score. Across five LLMs, we identify a six-layer residual evidence band whose state causally shifts the competition between attacker and correct answers. The same functional interface appears across models, while its downstream readout varies by architecture. Some models rely partly on a small set of attention heads, whereas others use more distributed pathways. Source and position controls show that the band carries answer-bearing document information rather than a generic signal that a document is present. On 106 answerable attacks, strong oracle steering suppresses the attacker answer in 92.5% of cases but recovers the correct answer in only 13.2%. Restoring clean residual states within the six-layer band recovers 71.7%, nearly matching the 72.6% recovery from restoring all layers. These results show that suppressing poisoned evidence is not sufficient for correction. Correct recovery requires restoring positive support for the right answer, and the residual band provides a compact causal interface for doing so.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.