acceptodds
Under review as a conference paper at ICLR 2027

Localize to Detect: Low-Budget Syntactic Span Watermarking for LLMs

Abstract

LLM watermarking provides a practical mechanism for identifying machine-generated text. Conventional methods perturb token probabilities at every decoding step, allowing these interventions to accumulate and degrade generation quality. Sparse watermarking addresses this issue by restricting watermarking to selected tokens. However, existing sparse statistical methods remain token-level, with no shared structural unit for embedding and detection. We introduce SpanWM, a span-based watermarking framework that uses parser-recoverable phrases as the unit of both embedding and detection. A secret key selects syntactic phrases for watermarking, and the detector re-identifies the corresponding spans from the received text for localized detection. This span-level design concentrates watermark evidence within a low token budget while leaving most of the output untouched. Experiments show that existing sparse methods must test substantially more tokens to match SpanWM's detection performance. SpanWM also better preserves the draft's meaning and remains robust to paraphrasing and token-level edits. At matched AUROC on C4, SpanWM requires only 16–45% of the watermark budget needed by the sparse baselines. Our code is available at https://anonymous.4open.science/r/spanwm_11-BCD6.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.