acceptodds
Under review as a conference paper at ICLR 2027

Tree-Based Speculative Decoding for Watermark Generation via Vocabulary Partitioning

Abstract

Watermarking helps identify LLM-generated text for data provenance, but its practical deployment requires fast generation. Tree-based speculative decoding could accelerate watermark generation by proposing multiple candidate continuations at each step, so that another candidate can still be accepted if one is rejected. However, applying it to watermarking is challenging: different branches use different pseudorandom sources, making the watermark signal harder to detect. We address this challenge by deterministically partitioning the vocabulary into disjoint subsets and sampling one draft token from each subset to form a distinct branch of the draft tree. The subset containing an observed token identifies the only draft branch that could have produced it, leaving that branch and residual sampling as the two possible sources. We prove that the resulting algorithm preserves the target model’s output distribution and achieves a draft acceptance probability at least as high as that of single-sequence speculative decoding. Experiments with Gumbel-max and SynthID watermarks show improved generation throughput while largely preserving detection power. We also show that our method can benefit from advances in drafter design, with an EAGLE-style draft head enabling additional speedups.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.