Exploring Unified Latent Diffusion for Multimodal Modeling
Abstract
Unified multimodal models commonly combine autoregressive language generation with continuous image diffusion, requiring a shared model to learn two different generative processes. We study Unified Latent Diffusion, which generates both language and images through continuous latent denoising, and ask how this choice affects joint multimodal learning. Through controlled pretraining from scratch, we compare autoregressive, discrete diffusion, and latent diffusion language modeling with matched backbone architectures, training data, and token budgets, while keeping visual modeling fixed. We identify three findings: (i) Unifying generative modeling through latent diffusion eases joint training, surpassing the autoregressive baseline's visual generation (DPG-Bench) with one-third fewer training tokens. (ii) Iterative denoising refines answers and strengthens visual understanding (Vision-Centric: 33.76 vs. 20.84 for AR at 30B training tokens). (iii) Effective language latents for diffusion modeling require different properties from vision. These comparisons examine the capabilities and limitations of unified latent diffusion as a generative formulation for general-purpose UMMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.