Poisoning by Dilution: Amplifying RAG Attacks with Answer-Free Documents
Abstract
Retrieval-augmented generation (RAG) is vulnerable to corpus poisoning, where injected passages steer a model toward an attacker-chosen incorrect answer. Existing attacks primarily construct passages that directly support this answer. We investigate whether passages without answer support can also strengthen targeted poisoning. We propose evidence dilution, which combines target-bearing passages with query-relevant, answer-free passages that compete for limited retrieval slots and alter the evidence available to the generator. Our method uses relation-aware generation, semantic checks, and proxy retrieval scores to construct and select dilution passages. We evaluate evidence dilution on Natural Questions, HotpotQA, and MS MARCO, achieving strict attack success rates of 92%, 87%, and 75%, respectively. On Natural Questions, adding three dilution passages increases strict attack success from 84% to 92% in vanilla RAG and from 77% to 91% under InstructRAG-ICL. Evidence analysis links these gains to changes in retrieved evidence and transitions from mixed answers to exclusive adoption of the attacker-chosen answer. These results show that passages can strengthen targeted corpus poisoning without directly supporting the attacker's answer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.