acceptodds
Under review as a conference paper at ICLR 2027

Backdooring Masked Diffusion Language Models

Abstract

Masked diffusion language models (MDLMs) are emerging as a compelling new paradigm for text generation, but their training-time security remains largely unexplored. MDLMs are trained with discrete-state corruption and generate text through iterative denoising. In this work, we present a systematic study of training-time backdoor attacks on MDLMs. We propose SHADOWMASK, a backdoor attack that modifies the forward corruption process on poisoned examples by replacing the standard all-mask terminal distribution with a trigger–mask mixture prior. Clean examples use ordinary mask-based training. We further provide a principled mathematical formulation, deriving a weighted denoising cross-entropy objective for this construction. Deployment uses the unmodified sampling procedure provided with each model. Evaluations on DiT-based MDLMs and four large diffusion language model families across WikiText-103, OpenWebText, and Alpaca show that SHADOWMASK achieves near-100% attack success in many settings, substantially outperforms the evaluated data-poisoning baseline in ASR, and largely preserves clean utility. It remains effective under full-model training and parameter-efficient fine-tuning, and persists under the evaluated defenses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.