acceptodds
Under review as a conference paper at ICLR 2027

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

Abstract

Continuous latent-space modeling has proved effective for image, video, and audio generation, while text generation is still dominated by models that directly predict discrete tokens. Existing continuous language models often inherit embedding spaces not designed for both generation and decoding, or compress autoencoded latents to ease diffusion, potentially sacrificing token-level fidelity. Instead of simplifying the representation to accommodate the generative model, we introduce AURORA-LM, a continuous-latent diffusion language model that preserves a high-capacity, decodable text latent and learns its distribution separately. A Query-based Encoder-Decoder organizes text into a prefix-aligned latent sequence. We then freeze the autoencoder and train a Block-causal Diffusion Transformer to model these full-width latents through blockwise flow matching. However, retaining these high-dimensional latents makes diffusion modeling more challenging. We address this difficulty by restricting only the noisy-input pathway through a low-rank bottleneck while retaining the full clean-latent prediction target, and calibrating the noise-level distribution to the latent width. Finally, we introduce self-trajectory consistency to bridge the gap between training on independently sampled noisy states and inference through iterative denoising. AURORA-LM achieves the best Gen-PPL and MAUVE on OpenWebText free generation and the best ROUGE scores on XSum conditional summarization among the evaluated continuous and diffusion-based language models. Scaling its denoiser to 1B parameters, it also outperforms a larger publicly released latent-diffusion language model under matched prompting and scoring.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.