Multi-Bit Watermarking for Discrete Diffusion Language Models via Bitwise Evidence Feedback
Abstract
Discrete diffusion language models (dLLMs) offer a new paradigm for language generation, but reliable multi-bit watermark embedding and recovery remain challenging. Iterative denoising involves partially masked, non-causal contexts and uneven evidence accumulation across message bits. This work focuses on embedding multi-bit messages during dLLM generation and recovering them from the resulting text. We propose BEFMark, a watermarking framework that uses bitwise evidence feedback to adapt embedding throughout denoising. BEFMark extracts bitwise evidence from partially masked sequences through context-aware multi-layer voting and estimates candidate tokens’ expected evidence contributions under incomplete contexts. This feedback guides a KL-regularized update of the generation distribution, strengthening weakly supported bits while penalizing deviation from the base model’s predictions. Experiments on LLaDA-8B-Instruct and Dream-v0-Instruct-7B across multiple datasets demonstrate improved bit recovery accuracy, with additional exact-message recovery gains on LLaDA-8B-Instruct. BEFMark also improves recovery robustness under text editing while maintaining competitive log-perplexity and downstream task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.