Tree-Based Speculative Decoding for Watermark Generation via Vocabulary Partitioning
Abstract
Watermarking helps identify LLM-generated text for data provenance, but its practical deployment requires fast generation. Tree-based speculative decoding could accelerate watermark generation by proposing multiple candidate continuations at each step, so that another candidate can still be accepted if one is rejected. However, applying it to watermarking is challenging: different branches use different pseudorandom sources, making the watermark signal harder to detect. We address this challenge by deterministically partitioning the vocabulary into disjoint subsets and sampling one draft token from each subset to form a distinct branch of the draft tree. The subset containing an observed token identifies the only draft branch that could have produced it, leaving that branch and residual sampling as the two possible sources. We prove that the resulting algorithm preserves the target model’s output distribution and achieves a draft acceptance probability at least as high as that of single-sequence speculative decoding. Experiments with Gumbel-max and SynthID watermarks show improved generation throughput while largely preserving detection power. We also show that our method can benefit from advances in drafter design, with an EAGLE-style draft head enabling additional speedups.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.