acceptodds
Under review as a conference paper at ICLR 2027

Watermarking Continuous Diffusion Language Models

Abstract

Existing watermarking methods for text generation focus on autoregressive and discrete diffusion language models, leaving continuous diffusion language models (CDLMs) unexplored. To close this gap, we introduce COMPASS (Continuous-Latent Output Marking for Provenance via Anchor Steering and Scoring), a training-free watermark designed specifically for CDLMs. COMPASS constructs key-dependent target and non-target anchor representations in the latent diffusion space and, during a selected denoising window, steers latent vectors associated with non-target anchors toward the target anchor set. Detection requires only the candidate text. The detector re-encodes the text using the CDLM’s encoder and measures whether the resulting latent representations align with the target anchors defined by the watermark key. We show that this design provides robust detection of content generated by state-of-the-art CDLMs, namely Cola and ELF-L, while preserving their high-quality outputs. Furthermore, COMPASS requires neither model fine-tuning nor additional neural-network evaluations during generation, providing an efficient and practical framework for provenance tracing in CDLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.