A Unified Framework for Token-Level Watermarks
Abstract
LLM watermarks allow tracing AI-generated texts by inserting a detectable signal into their generated content. Recent works have proposed a wide range of token-level watermarking algorithms, each with distinct designs. Crucially, there is currently no general and principled formulation for token-level LLM watermarking that could enable deeper understanding of current schemes and ease the design of new variants. In this work, we propose the first unified framework for token-level watermarks based on a principled constrained optimization problem. Our formulation unifies most existing and popular token-level watermarking methods, and explicitly reveals the constraints that each method optimizes. In the process, it highlights an underexplored quality-diversity-power trade-off. Our experimental evaluation validates our framework: watermarking schemes derived from a given constraint consistently maximize detection power with respect to that constraint. At the same time, the framework provides a principled approach for designing novel watermarking schemes tailored to specific requirements, for instance minimizing perplexity of the watermarked text.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.