acceptodds
Under review as a conference paper at ICLR 2027

Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning

Abstract

Multimodal models increasingly think with different modalities such as images, 3D point clouds, and robot states, not just text. Yet each modality is still encoded into its own representation space, so every step that switches modality must do two things: reason one step further, and translate between the formats of the two modalities. We call this cost the modality-switching gap. In this paper, we introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that unifies different modalities into a shared latent space for multimodal reasoning. A unified encoder maps teacher reasoning steps from different modalities into latent thought tokens in a shared space, trained to preserve the information needed for later reasoning steps and the final output. A diffusion reasoner, trained jointly with the encoder, generates these tokens at inference without teacher observations. Across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, Uni-LaDiR achieves relative gains over the strongest baselines of 7.3% on four mathematical and logical VLM benchmarks and 6.1% on RLBench manipulation tasks. Controlled comparisons show increasing gains as more teacher modalities are unified. These results suggest that unification improves multimodal reasoning by weaving it into a single thread, where the model advances from thought to thought without translating between formats.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.