Dataset Watermarking with Provable Black-Box Detection
Abstract
LLMs are often pre-trained and post-trained on vast amounts of loosely curated data, and potentially have been trained on proprietary datasets or the benchmarks used for evaluation. This motivates dataset watermarking: designing datasets such that training on them leaves detectable signatures in the resulting model. Existing approaches embed watermarks by either injecting artificial content or paraphrasing full examples under token-level generation control, and many require model access beyond generated outputs for detection. Our method embeds a dataset-level watermark by increasing the co-occurrence of randomly selected word pairs through meaning-preserving local edits, and detects it from generated text alone with provable false positive control. We conduct evaluation on four base models and three datasets and show that it reliably detects the watermark () in the fine-tuning stage, outperforming existing black-box baselines. Our method remains effective under substantial dilution where watermarked data is less than 5% of fine-tuning tokens. Moreover, our method better preserves benchmark utility and semantic integrity than existing baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.