Breaking the Factorization Barrier in Diffusion Language Models with External Memory
Abstract
Masked diffusion language models generate tokens in parallel, but factorized reverse steps cannot represent dependencies among simultaneously decoded tokens. At a fixed conditioning state, the KL divergence from the true joint conditional to the product of its exact marginals equals the block's total correlation. We introduce R-Joint, a training-free decoding framework that injects joint token structure into frozen diffusion language models through an editable external memory. Retrieved clean text spans define an empirical joint distribution over masked token blocks. We combine this distribution with the model's factorized proposal through a closed-form KL barycenter, with exact normalization using the sparse retrieval support. A gate tuned on validation text selects blocks for joint commitment. Experiments on LLaDA-8B and Dream-7B across Wikipedia, Python code, legal opinions, and synthetic entity corpora show consistent improvements in block exact match and likelihood at block length two. R-Joint improves exact match by up to percentage points on global real-world corpora and by up to points with repository-local memory. On Dream-7B, gains reach points on synthetic entity spans and points with denser memory. Matched ablations at the same block length show that preserving joint structure contributes beyond token-wise retrieval. At identical commit decisions, replacing the fused distribution with the product of its marginals reduces block exact match by to points. In iterative restoration of partially masked windows, joint commitment improves token accuracy with matching memory, even against a base decoder given more steps, and complements retrieved context. These results highlight the representational limits of factorized reverse steps and show that editable external memory can supply joint structure without retraining the model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.