acceptodds
Under review as a conference paper at ICLR 2027

Towards Statistically Optimal and Semantically Robust Watermarking for Large Language Models

Abstract

The widespread adoption of large language models (LLMs) in real-world applications poses critical challenges for reliably detecting LLM-generated text. Logit-based watermarking is a prominent approach that enables low-overhead statistical detection by biasing token selection through keyed binary masks. However, existing methods often suffer from limited detection signals under constrained perturbations, and remain vulnerable to content-preserving attacks. To address the detectability limitation, we theoretically establish a first-order characterization of watermark detectability in terms of the predictive probability mass induced by the vocabulary partition and derive a closed-form optimality condition that maximizes the expected detection -score under a fixed perturbation strength. Building on this result, we propose SoptiMark, which combines a Cumulative Probability Mask that approximates this optimal allocation with an Embedding Clustering mechanism that stabilizes mask reconstruction across semantically related tokens. This design strengthens the standardized watermark signal while improving resilience to content-preserving modifications. Extensive experiments demonstrate that SoptiMark consistently outperforms prior logit-based watermarking methods by preserving substantially stronger detection signals under diverse content-preserving attacks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.