Scaling and Distilling Text Embeddings for Better Diffusibility
Abstract
Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that *scaling* up to the advanced and general-purpose T5Gemma-2 embeddings greatly improves generative performance. But such raw embeddings are not optimal, as they are overly discriminative and even the embeddings of plausible alternative words are separated. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding in between. To address this, we *distill* T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves **Gen. PPL 17.8** (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M as an AR counterpart.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.