VeriRAG: Verifying Answer Updates for Robust RAG
Abstract
In retrieval-augmented question answering, a system may revise its retrieval-free initial answer using retrieved passages. Such updates can correct model errors, but corpus poisoning and indirect prompt injection can also make them harmful. A verifier can guide this decision by checking contextual support, yet may continue to endorse a candidate even after necessary support is removed. We introduce VeriRAG, a training-free, attack-label-free method that requires two complementary verifier responses before adopting a candidate: continued endorsement when complete support is preserved, and an insufficient-evidence judgment when necessary support is withdrawn. For questions whose answers require a finite set of factual relations to hold jointly, VeriRAG derives these relations from the question alone and selects the candidate to test from the candidate answer strings, without showing the retrieved passages to this selection step. Holding the question and candidate fixed, it constructs paired contexts for each relation: one preserves complete support, while the other removes all support for that relation but retains joint support for the others. All pairs are audited and frozen before response testing. Adoption requires the expected responses on every pair, complete and unopposed support in the original context, valid citations, and no fully supported competing answer. Any failed or unresolved check leaves the initial answer unchanged. Through this process, VeriRAG allows retrieval to correct initial errors while reducing harmful answer updates. Experiments across three datasets and three language models demonstrate improved robustness under mixed corpus-poisoning and indirect prompt-injection attacks, with an average end-to-end accuracy of 71.0% and an attack success rate of 5.6%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.