acceptodds
Under review as a conference paper at ICLR 2027

Implicit Ranking Leakage in Black-Box RAG: Retrieval Stability-Guided Knowledge Corruption Attacks

Abstract

Retrieval-Augmented Generation (RAG) is widely used in large language model systems. Recent studies show that RAG is vulnerable to knowledge corruption attacks (KCAs), where adversaries observe top-k retrieved documents to approximate the retriever and manipulate downstream generation. Existing attacks mainly use top-k membership as binary supervision and treat all retrieved documents as equally important. This ignores hidden ranking information and underestimates the risk exposed by retrieval behavior. In this paper, we propose Ranking Vulnerability-enhanced Knowledge Corruption Attack (RV-KCA), a black-box method that infers relative document retrieval stability from top- membership and uses it as pairwise supervision to conduct stronger knowledge corruption attacks. First, RV-KCA applies semantics-preserving query perturbations and estimates retrieval stability from how often each document remains in the top- set. Our theoretical analysis shows a monotonic relationship between retention probability and the original retrieval score under a shared local noise model. It also characterizes sampling uncertainty in the stability estimates, which serve as noisy evidence of relative document ordering. Second, RV-KCA selects document pairs with clear stability differences and uses them to train a margin-aware surrogate retriever. The surrogate model learns to assign higher scores to more stable documents while preserving membership supervision information derived from the observed top- sets. Finally, RV-KCA optimizes query-side triggers on the surrogate to give the injected document a ranking advantage over estimated boundary competitors across query variations. Experiments show that RV-KCA achieves higher surrogate fidelity and overall attack success rates than existing black-box baselines. It also reaches a fixed surrogate fidelity target with fewer target-retriever calls. These results demonstrate that retrieval stability exposes useful ranking information for knowledge corruption attacks even when similarity scores and explicit ranks are hidden.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.