Resilient RAG Dataset Inference via Orthogonal Dual-Layer Watermarking
Abstract
Retrieval-Augmented Generation (RAG) allows unauthorized service providers to exploit proprietary text corpora while hiding provenance behind neural paraphrasing and black-box access, making dataset theft difficult to prove. Lexical watermarks wash away under deep rewriting; injected semantic facts survive paraphrasing but are easily pruned by content filters, or are prohibited in regulated domains such as medicine and law where facts cannot be altered. The Thief's Dilemma formalizes a carrier-conditional bottleneck: under joint carrier contraction, suppressing both channels violates any utility requirement above the carrier-free factual rate. Fact-only reconstruction lies outside this family and can erase regulated-mode provenance. We present orthogonal dual-layer watermarking with a black-box Prober-Verifier. The semantic layer embeds mode-specific invariants: open-domain documents receive long-tail predicate bindings, whereas regulated documents receive truth-preserving coordinate permutations that leave the causal claim graph strictly intact. With semantics frozen, Bidirectional Masked Contextual Resampling (BMCR) introduces tokenizer-independent functional-word biases into disjoint syntactic slots. An active prober uses at most 40 target-side API calls, fusing complementary evidence via a directional Gaussian-copula test. Across a 150-setting sweep of carrier-retaining attacks on the RDI benchmark, our joint auditor maintains at least 98.17% detection power at a 0.1% false-positive rate. Under an aligned 40-call budget, dual-layer attribution reaches 99.35% power against adaptive mixed attacks, exceeding the active semantic channel alone by 15.27 percentage points. Finally, we provide finite text-edit certificates, prove noninferiority on downstream clinical and legal tasks, and show that amortized calibration reduces attribution overhead to 110.68 calls per audit in production monitoring. Code is available at https://anonymous.4open.science/r/rpd_anonymous-F278.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.