Latent Chains for Robust and Tamper-Evident Watermark of Diffusion Large Language Models
Abstract
Diffusion large language models (dLLMs) are emerging as an alternative to autoregressive LLMs, with blockwise generation becoming a prominent design. LLM watermarking provides a promising approach for identifying dLLM-generated text. However, existing dLLM watermarks achieve detectability and robustness but provide limited tamper evidence, leaving them vulnerable to piggyback spoofing attacks that modify watermarked text while preserving its attribution to the targeted model. To provide tamper evidence, we introduce Latent Chain Watermark (LCW), which uses a fixed latent chain within each block and lets each finalized block select the latent chain for the next block. To overcome the tension between robustness and tamper evidence, one favours unaffected by modification, the other favours sensitive to modification. We adopt two detections: (1) robust watermark detection searches over all candidate changes to recover watermark evidence, (2) tamper detection measures how modifications disrupt block boundaries and latent chains as tamper evidence. We further develop efficient methods to accelerate detection. At a 5% false-positive rate, LCW achieves 95.1% watermark detection, 80.2% watermark survival under 10% token-level editing, and 82.2% tamper detection after three token edits. Additional experiments demonstrate that LCW's applicability beyond diffusion generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.