acceptodds
Under review as a conference paper at ICLR 2027

Exploring Unified Latent Diffusion for Multimodal Modeling

Abstract

Unified multimodal models commonly combine autoregressive language generation with continuous image diffusion, requiring a shared model to learn two different generative processes. We study Unified Latent Diffusion, which generates both language and images through continuous latent denoising, and ask how this choice affects joint multimodal learning. Through controlled pretraining from scratch, we compare autoregressive, discrete diffusion, and latent diffusion language modeling with matched backbone architectures, training data, and token budgets, while keeping visual modeling fixed. We identify three findings: (i) Unifying generative modeling through latent diffusion eases joint training, surpassing the autoregressive baseline's visual generation (DPG-Bench) with one-third fewer training tokens. (ii) Iterative denoising refines answers and strengthens visual understanding (Vision-Centric: 33.76 vs. 20.84 for AR at 30B training tokens). (iii) Effective language latents for diffusion modeling require different properties from vision. These comparisons examine the capabilities and limitations of unified latent diffusion as a generative formulation for general-purpose UMMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.