acceptodds
Under review as a conference paper at ICLR 2027

Reconciling privacy with provenance for language model watermarks

Abstract

Watermarking has great potential for regulating the use of language models (LMs), but its deployment has been marred by controversy, as many users argue that fine-grained tracking of LM usage violates their privacy. This tension reduces the effectiveness of watermarks, as users engage in evasion strategies such as paraphrasing that can remove much of the watermark signal. We show that provenance and privacy can be reconciled if we shift the goal of watermark design from tracking individuals to estimating aggregate usage. We develop locally differentially private watermarks which provably limit the evidence that any individual text can contribute to detection, while still enabling accurate estimation of population frequencies. We introduce the Red-Green Sequence (RGS) watermark, a quality-preserving and computationally efficient watermark that achieves accurate frequency estimation at high privacy levels across various models, sample sizes and generation entropies in settings that reflect realistic deployment scenarios. On C4 language modeling tasks, we attain MSE with bias at a low of 1, while in more challenging settings like an essay writing task with very few text samples, we still find errors on the order of MSE with bias at .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.