CleanBase: Detecting Multi-Document Prompt Injection in RAG Databases
Abstract
Retrieval-augmented generation (RAG) systems are vulnerable to prompt injection attacks that insert malicious documents into their knowledge databases. Many existing RAG defenses intervene only after the database has already been contaminated, for example by hardening the language model or filtering retrieved content at query time. These approaches have complementary but distinct limitations: model-level defenses are tied to the serving model, while query-time defenses repeatedly incur security processing during inference; moreover, neither directly removes the underlying contamination from the knowledge database. We therefore explore a complementary perspective: treating the knowledge database itself as a proactive security layer and sanitizing it before retrieval and generation. We introduce CleanBase, a training-free offline sanitizer that exploits cross-document semantic relationships in the knowledge database. CleanBase constructs a -nearest-neighbor graph over documents, prunes weak edges using a data-adaptive threshold, and removes documents participating in tightly connected groups. It requires neither future user queries nor access to the serving language model, and introduces no additional CleanBase computation per query after sanitization. Across six knowledge databases and nine poisoning attacks, CleanBase reduces mean attack success rate from 69.9% to 9.1% at the evaluated operating settings while largely preserving aggregate RAG utility. We further characterize its limitations under low-cardinality and detector-aware diversified attacks. These results demonstrate the promise of database-level sanitization as a complementary defense layer for RAG systems. Our source code is available at https://anonymous.4open.science/r/CleanBase-7762.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.