acceptodds
Under review as a conference paper at ICLR 2027

PoisonSieve: Detecting Poisoned Documents in RAG through Matched Local Contrast

Abstract

Retrieval-augmented generation uses an external corpus as inference-time evidence, allowing an attacker to promote a false answer by injecting a handful of documents. Detection must distinguish this manipulation from ordinary relevance without knowing which queries or documents are targeted. Existing detectors use text irregularity, candidate consensus, or corpus-level graph structure, whose reliability varies with the attack and local context. We present PoisonSieve, which constructs a reference matched to each detection scope. At query time, PoisonSieve-Query (PSQ) compares generation candidates with the lower-ranked tail of the same retrieval, exposing answer-token concentration and carrier–payload seams. At corpus time, PoisonSieve-Graph (PSG) compares each document's strongest semantic relations with its own neighborhood floor to measure coordinated density, complemented by script integrity. Neither requires poison labels, a trusted corpus, or training. Across three QA datasets, three dense retrievers, and six poisoning constructions, PSQ reaches 95.2% AUROC and detects 82.2% of poison at a 5% clean-removal budget, against 81.1% and 52.5% for the strongest query-time baseline; PSG reaches 93.3% and 79.8% against 79.4% and 37.6% for the strongest corpus-time baseline, with a 79.6% versus 1.4% detection rate on camouflaged injections. Joint deployment cuts attack success from 67.4% to 16.1% while retaining unpoisoned-retrieval F1 at 41.0%, compared with 42.1% without filtering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.