SIDRAG: Adaptive Defense against RAG Poisoning via Set-Level Influence Detection
Abstract
Retrieval-augmented generation (RAG) systems are highly vulnerable to poisoning attacks, and existing defenses struggle to maintain robustness across varying attack scales. We observe that successful attacks typically rely on malicious documents exerting either an overwhelmingly strong individual influence on the large language model's (LLM) predictions, or a consistent and strong aggregated influence when acting jointly. Motivated by this, we propose SIDRAG, a defense framework that adaptively mitigates poisoning attacks across different scales by analyzing document influence patterns. Specifically, SIDRAG decomposes the LLM's attention output into document-level influence vectors and constructs an attributed signed graph to capture their strengths and pairwise directional relationships. It then identifies Jointly Consistent Subsets (JCSs) within the graph, grouping documents with aligned influence patterns. This set-level modeling unifies the coordinated impact of multiple documents into a single detection unit, so that the defense does not require prior knowledge of the number of malicious documents. Finally, SIDRAG flags anomalous subsets based on their influence strength and consistency, filtering out malicious documents prior to answer generation. Extensive experiments demonstrate that SIDRAG achieves a 98.7% average exact identification rate for malicious document sets. Across five LLMs and three benchmarks, it achieves the highest average robust accuracy (RACC) and the lowest average attack success rate (ASR) among existing defenses, while incurring the lowest inference latency among defenses on every evaluated model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.