Designing Transformers for Continuous Diffusion Language Models
Abstract
What should a Transformer architecture preserve, learn, and feed back when language is generated by continuous denoising? In embedding-space diffusion, the denoiser connects a continuous corruption process to categorical prediction, while its outputs determine both state updates and auxiliary inputs. We study architecture through these distinct computational roles and formulate three design principles: separate feature-scale control from noise-dependent modulation; separate the forward weighting of geometric token evidence from its gradients to the codebook; and distinguish clean-state estimation from self-conditioning. We instantiate these principles in a LangFlow-style model through normalization, shortcut gradient routing and calibration, and noise and confidence conditioning. Elementary analyses make the scope of the design explicit: normalized attention limits dependence on raw query–key magnitudes, the Gaussian observation model explains the analytic shortcut, and shrinking a posterior mean need not improve clean-state estimation. The architecture retains the baseline representation, objective, and inference sampler. A seven-stage OpenWebText-1024 ablation provides the core experimental test, alongside comparisons with FLM, RePlaID, LangFlow, and ELF implementations. Mechanism controls and transfer evaluations distinguish improvements to one implementation from evidence for a reusable CDLM architecture.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.