acceptodds
Under review as a conference paper at ICLR 2027

PRESENCE BEFORE PAYLOAD: FALSE-POSITIVE- CONTROLLED MULTI-BIT LLM WATERMARKING

Abstract

Multi-bit LLM watermarking involves two conceptually distinct tasks: detecting whether a watermark is present and, only if so, recovering the payload it encodes. In several existing methods, however, the two decisions are conflated: watermark presence is inferred from the same search-and-decoding procedure used to recover the payload. This can cause unwatermarked text to be accepted whenever payload search or decoding produces a plausible candidate, making nominal false-positive rate (FPR) guarantees unreliable, especially when the assumed null distribution does not match real unwatermarked text. We introduce PIPER, a presence-before-payload watermark that separates the two decisions by design. PIPER partitions the vocabulary into two keyed levels so that the same token observations support both decisions: a coarse level provides payload-independent evidence for presence testing, while a finer level carries payload information for recovery. Presence is decided by an exact binomial test over deduplicated effective contexts, and the payload is decoded only if this test accepts. Under the standard random-key model with an idealized PRF, we prove finite-sample FPR control for a presence test whose statistic and threshold are independent of payload size and decoder choice. We evaluate PIPER across generation lengths from \(T=100\) to \(T=1000\). At the representative \(T=300\) setting and a nominal \(1%\) FPR, PIPER requires no held-out threshold fitting under the stated null model and achieves empirical FPRs of \(0.86%\) and \(0.92%\) on model-generated and natural-text nulls, respectively, with \(98.7%\)–\(99.0%\) presence TPR across 8- and 16-bit payloads.Code is available at [this URL](https://anonymous.4open.science/r/PIPER-F6B3/).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.