Joint Discrete Diffusion
Abstract
Discrete diffusion models generate tokens in any order, making conditionally independent predictions within each denoising step. Sampling multiple tokens independently in parallel can however severely degrade generation quality, motivating tractable models of joint token distributions. Here we introduce joint discrete diffusion, which uses a small learned conditional random field (CRF) head on a frozen diffusion backbone. The head defines a joint distribution over the sequence at each denoising step, allowing exact normalization, marginalization, and joint sampling. Two novel contributions are introduced to make this idea work in practice: A low-rank parameterization of the CRF that allows efficient use of the conditioning sequence and a one-step reveal objective. This reveal the remaining tokens objective is motivated by the fact that the CRF ELBO objective in the low reveal rate regime assigns a very low probability to revealing adjacent tokens simultaneously. Supervision of token dependence is therefore by construction substantially weakened. We tested our approach across large diffusion language models iLLaDA, Dream, SDLLM-MDLM, and SDLLM-Duo in prompt continuation tasks. Comparing the joint against the native backbone model, we observe over a range of conditions that the native is outperformed by a joint with half the number of generation steps. Our joint token modeling achieves a measured 1.8–1.9x speedup at equal or better scores under both text evaluators.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.