acceptodds
Under review as a conference paper at ICLR 2027

AF-RAG: Active Falsification for Local Top-K Contamination in Literature-Grounded Question Answering

Abstract

Automated laboratories and auto-research systems increasingly use retrieved literature to generate hypotheses, plan experiments, and analyze results. However, erroneous numerical values, conclusions, or comparisons in the retrieval context can propagate into downstream research decisions. We study a constrained local retrieval failure in which erroneous or adversarially promoted evidence occupies the initial Top-*K*, while claim-relevant corrective evidence remains retrievable at deeper ranks in the same corpus. We propose Active Falsification RAG (AF-RAG), which audits verifiable claims, generates targeted falsification queries, retrieves counter-evidence beyond Top-*K*, and adjudicates claim-level conflicts. We formalize this condition as an information cocoon with stealth poisoning, where locally coherent micro-tampering changes query-relevant facts. Under this local assumption, AF-RAG conditionally prioritizes claim-relevant deep evidence. On 2,641 arXiv papers and 1,200 stealth-poisoned samples, AF-RAG lowers Poisoned Adoption Rate (PAR) from 94.75% to 42.67%, raises Truth Recovery Rate (TRR) from 25.58% to 67.50%, increases Joint Success Rate (JSR) from 3.75% to 49.83%, and achieves 79.9% accuracy on unpoisoned queries. Supplementary diagnostics evaluate a verification extension combining provenance-aware adjudication and conflict-driven evidence acquisition, showing improved joint success when initial evidence is correct and external evidence is contaminated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.