acceptodds
Under review as a conference paper at ICLR 2027

Is Binary Enough? Exploring P-Ary Encoding for Multi-Bit Language Model Watermarking

Abstract

Large language model (LLM) watermarking provides a practical mechanism for tracing AI-generated text and mitigating misuse. Recent multi-bit watermarking methods enable richer payloads, yet almost all rely on binary encoding, leaving the role of the encoding radix largely unexplored. We ask a fundamental question: can a higher-radix representation outperform binary encoding for multi-bit watermarking? We introduce PAMark, a lightweight watermarking scheme that combines \(p\)-ary encoding with a Gumbel-based ranking strategy. Through theoretical analysis, we characterize how the encoding radix affects message recovery and prove that, under certain conditions, an appropriate choice of \(p\) achieves higher message accuracy than binary encoding. Extensive experiments across multiple models and settings further validate our analysis. Compared with representative multi-bit watermarking baselines, PAMark achieves consistently stronger message recovery while maintaining competitive text quality and watermark detectability. Our results identify the encoding radix as an important yet overlooked design dimension and provide a principled foundation for moving beyond binary multi-bit LLM watermarking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.