AF-RAG: Active Falsification for Local Top-K Contamination in Literature-Grounded Question Answering
Abstract
Automated laboratories and auto-research systems increasingly use retrieved literature to generate hypotheses, plan experiments, and analyze results. However, erroneous numerical values, conclusions, or comparisons in the retrieval context can propagate into downstream research decisions. We study a constrained local retrieval failure in which erroneous or adversarially promoted evidence occupies the initial Top-*K*, while claim-relevant corrective evidence remains retrievable at deeper ranks in the same corpus. We propose Active Falsification RAG (AF-RAG), which audits verifiable claims, generates targeted falsification queries, retrieves counter-evidence beyond Top-*K*, and adjudicates claim-level conflicts. We formalize this condition as an information cocoon with stealth poisoning, where locally coherent micro-tampering changes query-relevant facts. Under this local assumption, AF-RAG conditionally prioritizes claim-relevant deep evidence. On 2,641 arXiv papers and 1,200 stealth-poisoned samples, AF-RAG lowers Poisoned Adoption Rate (PAR) from 94.75% to 42.67%, raises Truth Recovery Rate (TRR) from 25.58% to 67.50%, increases Joint Success Rate (JSR) from 3.75% to 49.83%, and achieves 79.9% accuracy on unpoisoned queries. Supplementary diagnostics evaluate a verification extension combining provenance-aware adjudication and conflict-driven evidence acquisition, showing improved joint success when initial evidence is correct and external evidence is contaminated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.