acceptodds
Under review as a conference paper at ICLR 2027

A Practical and Undetectable Watermark for Language Models Using Pseudorandom Codes

Abstract

The strongest quality notion for language model watermarks is undetectability, which ensures that the watermarked model is computationally indistinguishable from the original model. This notion is powerful, implying that the watermark preserves the model’s performance on any task that can be run in polynomial time. However, until now all undetectable watermarks either lacked nontrivial robustness (Christ et al., 2024) or had no practical implementation (Christ and Gunn, 2024). The latter construction uses pseudorandom codes (PRCs), whose codeword length corresponds to the number of tokens required for detection. Achieving nontrivial robustness required infeasibly long codewords, on the order of thousands of tokens. In this work, we significantly reduce the number of tokens required for detection in the PRC watermarking scheme of Christ and Gunn (2024) by introducing a new entropy-aware detector. We show that on the Qwen family of models, reliable detection is possible given only approximately 450 tokens. For the first time, this brings robust and undetectable LLM watermarks into the realm of practicality. Empirically, we find that the PRC watermark resists black-box detection and preserves quality, as measured by benchmark performance and output diversity. Compared with state-of-the-art baselines, our watermark offers greater resistance to the evaluated watermark-stealing attacks than SynthID-Text (Dathathri et al., 2024) and more diverse responses than TextSeal (Sander et al., 2026). Thus, we provide the first practical implementation of PRC watermarking for language models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.