acceptodds
Under review as a conference paper at ICLR 2027

Fused Multi-Latent Reasoning: Unified Multimodal Reasoning in Latent Space

Abstract

Vision-language reasoning often unfolds through intermediate steps that mix textual reasoning, visual evidence, and cross-modal grounding. Existing latent reasoning methods attempt to move such processes into latent space, but they typically rely on unimodal or weakly coupled supervision, leaving latent units poorly aligned with the multimodal reasoning steps they replace and prone to losing task-relevant visual evidence or semantic intent during compression. Our key insight is that high-dimensional vision-language embeddings can compress interleaved textual-visual reasoning sequences into dense representations that, after adaptation, serve as informative latent units for multimodal reasoning. In this work, we propose Fused Multi-Latent Reasoning (FMLR), a latent-space reasoning framework that learns latent units supervised by fused textual-visual step representations. By distilling step-level multimodal embeddings into latent units, FMLR enables the model to reason over implicit multimodal states. To further improve latent reasoning and control its length, we introduce Minimum-Sufficient Latent GRPO (MSL-GRPO), which combines answer correctness, threshold-based early sufficiency, and a per-unit latent cost to learn concise reasoning and explicit stopping decisions. Experiments on challenging vision-language reasoning benchmarks show that FMLR improves accuracy while producing substantially fewer output tokens than explicit-reasoning baselines. Code and datasets will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.