Adaptive Payload Allocation for Multi-bit LLM Text Watermarking
Abstract
Watermarking has emerged as a promising technique for tracing the provenance of content generated by large language models (LLMs). Compared to single-bit watermarking, multi-bit watermarking could embed richer provenance information, such as user identifiers. Most existing multi-bit methods adopt a fixed-payload design, embedding the same number of message bits at every generation step. However, we argue that the payload should be adaptively selected to maximize watermark detectability. In this paper, we characterize detectability using the mutual information since it quantifies the information about the embedded message after observing the sampled tokens. We further show that the achievable mutual information varies with payload size, which motivates formulating per-step payload selection as a mutual-information optimization problem. Based on this, we introduce Maximized Mutual-Information-Based Payload (MMIP), an oracle allocator that exhaustively evaluates candidate payload sizes and selects the one achieving the largest mutual information. However, MMIP requires constructing multiple message-conditioned watermarked distributions for each candidate payload, resulting in substantial inference overhead. To enable efficient online allocation, we further propose Entropy-Based Watermark Payload (EWP). Inspired by findings from MMIP, EWP derives the payload directly from the entropy of the original next-token distribution, thereby avoiding both payload enumeration and explicit construction of candidate watermarked distributions. MMIP and EWP are compatible with representative logit-based and distortion-free multi-bit watermarking methods. Experiments show that EWP achieves comparable performance with MMIP with negligible additional latency and both methods consistently improves the message-recovery–generation-quality trade-off over fixed-payload baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.