acceptodds
Under review as a conference paper at ICLR 2027

Structured Parallelism for Diffusion Language Models

Abstract

Discrete diffusion language models generate text by repeatedly updating noisy or masked tokens. At each step, a bidirectional Transformer backbone uses context from both directions to compute token representations, which an output layer converts into probabilities. Most models use a product distribution: they learn individual token predictions and sample them independently given the noisy sequence. This simplifies training and sampling, but even accurate individual predictions can yield inconsistent combinations. Later steps can improve consistency, at the cost of further Transformer passes. We introduce Structured-Parallel Diffusion Language Models (SPDLMs), which model dependencies among tokens proposed in the same step. At each denoising step, a lightweight decoder reuses a single Transformer evaluation of the noisy sequence to generate dependent tokens in parallel groups. Sequential computation is confined to this decoder, without reevaluating the Transformer within the step. Experiments on controlled pretraining and adaptation of existing diffusion language models show that this structured joint distribution improves generation with few denoising steps, reducing the need for costly backbone evaluations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.