Auditing Source Provenance in Retrieval-Augmented Data Synthesis
Abstract
Retrieval-augmented data synthesis transforms source documents into new datasets, making their provenance difficult to verify. We study whether document owners can audit such downstream use by modifying only their source documents, without controlling retrieval or generation or editing the synthesized outputs. We introduce a source-side semantic watermarking framework in which a secret key selects recommendations among alternatives with equivalent behavior for a given task. The key is encoded in recommendation relations, allowing the signal to be expressed without verbatim reproduction of source passages. Given access to a candidate derived dataset, a relation-aware auditor extracts expressed recommendations where available and aggregates group-level evidence to test alignment with the owner's key. We evaluate the framework under blind synthesis, where the generator is not instructed to preserve or reveal the watermark. Matched independent-key controls distinguish key-specific transfer from incidental generator preferences. Experiments on controlled technical-document benchmarks, including executable tasks and held-out content groups, demonstrate detectable source-specific signals after synthesis. Further analyses separate relation coverage from fidelity and characterize detection under partial source use, rewriting, and data mixing, as well as removal through informed filtering. These findings establish the feasibility of source-side provenance auditing for RAG-derived datasets and characterize the conditions under which its evidence is preserved or lost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.