acceptodds
Under review as a conference paper at ICLR 2027

GE: Grounded Exploitation–Exploration Extraction from Black-Box RAG Systems

Abstract

Retrieval-Augmented Generation (RAG) systems combine Large Language Models (LLMs) with external knowledge repositories, creating a new attack surface in which a proprietary corpus can be systematically extracted through conversational interaction. Explicit extraction attacks that directly request verbatim content are readily detected by lexical guardrails and refusal mechanisms, motivating stealthier semantic extraction through benign queries. Yet existing implicit probing methods remain vulnerable to semantic drift and localized exploration, limiting their ability to systematically cover the underlying corpus. In this work, we show that grounded extraction provides an effective design principle for addressing these limitations: generating candidate attack queries directly from returned responses mitigates out-of-domain semantic drift, while selecting attack queries via history-relative maximin distance optimization prevents extraction from becoming trapped within localized concept clusters. We further identify a complementary leakage channel arising from the discrepancy between retrieval and generation, which we term Echo Knowledge Overflow (EKO). Because production RAG systems often retrieve more passages than are needed to answer a query, portions of the retrieved context may remain unused in the response. However, this latent context can be surfaced as semantically adjacent follow-up cues through benign continuation-style instructions. Building on these observations, we introduce Grounded Exploitation–Exploration Extraction (GE), a black-box knowledge extraction attack that operates through benign natural-language queries. GE implements a dual-track strategy. Its exploration track derives progressive queries from facts confirmed in responses, while its exploitation track harvests EKO-derived cues using inductive continuation suffixes. A history-relative maximin distance algorithm then selects the next query from the response-grounded candidate pool to balance local exploitation with broader corpus exploration under a constrained query budget. Across four benchmark corpora and twelve victim LLMs, GE achieves macro-average Retrieval Coverage Rates of 0.847 on local models and 0.846 on enterprise endpoints within a 1,000-query budget, outperforming evaluated baselines by 15.1 and 21.3 percentage points, respectively, while sustaining a 96.8–97.6% Attack Success Rate against input-output guardrails. Beyond exposing retrieved knowledge, the extracted information enables the construction of Shadow RAG surrogates that replicate victim services with up to an 84.7% valid answer rate and 0.812 semantic answer similarity, demonstrating the practical feasibility of functional service replication.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.