Training Latent Language Diffusion Models Requires Less Than You Think
Abstract
Contextual latent diffusion offers a way to generate language in a representation space that captures linguistic dependencies beyond token identity, but existing approaches typically rely on pretrained language features or a separate representation-learning stage. We introduce LDLM-s: Scratch Latent Language Diffusion Model, to our knowledge, the first latent language diffusion framework to achieve strong generation while jointly learning its contextual representation space and generator from random initialization, without pretrained language components or a separate representation-pretraining stage. LDLM-s uses a language-modeling autoencoder to provide linguistic supervision to the latent space, while the diffusion objective simultaneously shapes the encoder, allowing the representation and denoiser to co-adapt throughout training. On OpenWebText, LDLM-s outperforms most of the recent continuous and discrete diffusion baselines on the generative-perplexity–token-entropy tradeoff. It also retains substantially more repetition-clean generations in the low-perplexity regime than the competing continuous baselines. Relative to LDLM, LDLM-s extends the tradeoff toward lower generative perplexity, while LDLM reaches higher-diversity operating points, despite LDLM-s seeing roughly fewer training tokens. More broadly, our results show that strong contextual latent diffusion can be learned without external language representations, providing a step toward self-contained latent language generation directly from text.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.