acceptodds
Under review as a conference paper at ICLR 2027

When Similarity Misleads: Query Perturbation for Robust Membership Inference in RAG

Abstract

Retrieval-Augmented Generation (RAG) systems enhance LLMs with external knowledge bases, but introduce new privacy risks. Membership inference attacks (MIAs) against RAG systems can exploit black-box access to infer whether a target document is contained in the knowledge base. Existing works typically rely on response accuracy, yet semantically similar non-member documents can also provide sufficient evidence for correct answers, causing false-positive inferences. To address this challenge, we propose QP-MIA, a black-box MIA based on query composition perturbation. QP-MIA constructs joint queries from target-specific fill-in-the-blank questions and perturbs their composition by replacing a subset of questions and reordering their positions, while preserving their connection to the target document. It then contrasts responses to two partially overlapping query compositions and extracts membership signals from changes in answer accuracy. This design distinguishes coherent evidence from genuine members from fragmented evidence provided by semantically similar non-members. Extensive experiments demonstrate an average AUC improvement of 20% over representative baselines across diverse RAG configurations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.