acceptodds
Under review as a conference paper at ICLR 2027

Joint Discrete Diffusion

Abstract

Discrete diffusion models generate tokens in any order, making conditionally independent predictions within each denoising step. Sampling multiple tokens independently in parallel can however severely degrade generation quality, motivating tractable models of joint token distributions. Here we introduce joint discrete diffusion, which uses a small learned conditional random field (CRF) head on a frozen diffusion backbone. The head defines a joint distribution over the sequence at each denoising step, allowing exact normalization, marginalization, and joint sampling. Two novel contributions are introduced to make this idea work in practice: A low-rank parameterization of the CRF that allows efficient use of the conditioning sequence and a one-step reveal objective. This reveal the remaining tokens objective is motivated by the fact that the CRF ELBO objective in the low reveal rate regime assigns a very low probability to revealing adjacent tokens simultaneously. Supervision of token dependence is therefore by construction substantially weakened. We tested our approach across large diffusion language models iLLaDA, Dream, SDLLM-MDLM, and SDLLM-Duo in prompt continuation tasks. Comparing the joint against the native backbone model, we observe over a range of conditions that the native is outperformed by a joint with half the number of generation steps. Our joint token modeling achieves a measured 1.8–1.9x speedup at equal or better scores under both text evaluators.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.