acceptodds
Under review as a conference paper at ICLR 2027

Dataset Watermarking with Provable Black-Box Detection

Abstract

LLMs are often pre-trained and post-trained on vast amounts of loosely curated data, and potentially have been trained on proprietary datasets or the benchmarks used for evaluation. This motivates dataset watermarking: designing datasets such that training on them leaves detectable signatures in the resulting model. Existing approaches embed watermarks by either injecting artificial content or paraphrasing full examples under token-level generation control, and many require model access beyond generated outputs for detection. Our method embeds a dataset-level watermark by increasing the co-occurrence of randomly selected word pairs through meaning-preserving local edits, and detects it from generated text alone with provable false positive control. We conduct evaluation on four base models and three datasets and show that it reliably detects the watermark () in the fine-tuning stage, outperforming existing black-box baselines. Our method remains effective under substantial dilution where watermarked data is less than 5% of fine-tuning tokens. Moreover, our method better preserves benchmark utility and semantic integrity than existing baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.