acceptodds
Under review as a conference paper at ICLR 2027

WOSpec: Watermarked Online Speculative Decoding

Abstract

Large language models (LLMs) increasingly rely on speculative decoding for efficient inference and on watermarking for output provenance. Online speculative decoding improves efficiency by training the draft model on the queries it serves, but when the watermark is embedded through the draft, every draft update can change the watermark carried by the generated text. Therefore, we show that to avoid this problem, drafter-invariant speculative decoding is needed, and in such an online setting the generated text is the watermarked text for every draft, including drafts trained on earlier watermarked outputs. In contrast, the watermark strength of watermarked speculative sampling with the standard acceptance rate depends on the draft, which was unstable. We propose WOSpec, a watermarked online speculative decoding framework built on watermarked Gumbel-max list sampling, which naturally has a drafter-invariance property, achieves desriable sampling efficiency, and maintains stable watermark detectability during online training. On Llama-2-7B-chat and Vicuna-13B, the watermark strength of our WOSpec remains unchanged under online training and matches that of a recently proposed watermark-maintaining sampling scheme, while WOSpec accepts 25–28% more tokens per verification step in terms of sampling efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.