One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion
Abstract
Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be dif- ficult to model with diffusion and decode reliably into tokens, which limits gen- eration quality after compression. To address this problem, we introduce JPEG- DLM (Joint-embedding Prediction for Efficient Generation with Diffusion Lan- guage Model), which jointly trains a compressor, a flow matching model and a de- coding module. With joint-embedding prediction, JPEG-DLM learns compressed embeddings that are more structured, easier to model with diffusion and reliably decodable into tokens. JPEG-DLM achieves the lowest mean Gen-PPL and high- est throughput among recent diffusion and flow models on LM1B and OWT. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 and approximately 2.3×ELF’s throughput. These results suggest that jointly learning compressed em- beddings offers a promising path toward efficient diffusion language modeling.We include our implementation code in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.