Fragile Watermark in Masked Diffusion Language Models
Abstract
With the rapid development of masked diffusion models, ensuring the integrity and security of their generated content has become increasingly important. Consequently, fragile watermarking has attracted considerable attention as a promising solution. It is designed to determine whether model outputs have been altered by third parties, thereby safeguarding their integrity. In contrast to robust watermarking, which is intended to remain detectable after common transformations such as translation and editing, fragile watermarking is invalidated by any unauthorized modification. It therefore functions similarly to peer-to-peer hashing, although the detector does not need to store a hash value. However, because masked diffusion models generate sequences with non-Markovian dependencies, autoregression-based fragile watermarking algorithms are difficult to apply, leaving this problem largely unexplored. To address this gap, we propose Dandelion: a block-dependent watermarking algorithm that verifies sequence integrity through the Markov states of individual blocks. We enable the minimum possible entropy increase at the watermark-embedding positions, thereby preserving the model’s generative capability. Experiments show that Dandelion achieves a 100% detection rate on unperturbed watermarked text. Under 13 types of perturbations, including substitution, truncation, and deletion, it achieves a watermark invalidation rate exceeding 99%. Our codes are available at: https://anonymous.4open.science/r/Dandelion-DLLM_Fragile_Watermark-32EF
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.