Semantic Diffusion Language Models via Low-Rank Corruption
Abstract
Diffusion language models are commonly trained to denoise text in which tokens have been replaced by a [MASK] or by tokens sampled uniformly from the vocabulary. Such replacements carry no information about the original token. To train on corruptions that more closely resemble plausible mistakes, we instead consider semantically related alternatives, which can be harder to identify as corrupted while preserving clues that aid recovery. We introduce Semantic Diffusion Language Models (SDLMs), a family of discrete diffusion models that draw these replacements from a nonnegative low-rank kernel, which permits exact training-loss evaluation at linear rather than quadratic cost in vocabulary size. Like uniform replacements, semantic ones are not marked as corrupted, so the model must learn not only what each token should become but also when to revise it. We show that with uniform noise, models time revisions by their confidence in the observed token alone: belief in a plausible alternative is indistinguishable from belief in an unrelated word. SDLMs instead use the full posterior, keeping a token settled when the model's belief falls on words similar to the observed token. SDLM exceeds the MAUVE of matched MDLM and GIDD baselines with 4× fewer steps. We also introduce SHARP, a predictor–corrector sampler that reallocates revisions toward tokens the model is most likely to change. With SHARP, SDLM outperforms ReMDM at every sampling budget and Duo with Ψ-samplers at all but the lowest.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.