acceptodds
Under review as a conference paper at ICLR 2027

MIRA : Multimodal Interleaved Reasoning with Adaptive Visual Tokens

Abstract

Vision-Language Models (VLMs) have demonstrated remarkable semantic understanding. Yet, their reliance on discrete linguistic tokens fundamentally bottlenecks performance on vision-centric tasks that demand fine-grained perceptual understanding, such as spatial localization, depth estimation, and detailed visual recognition. Recent advances in latent reasoning attempt to mitigate this by operating in continuous spaces. However, current latent methods either lack meaningful visual representations or adopt fixed reasoning templates that generate a predetermined set of visual tokens regardless of the input task. To overcome these limitations, we propose MIRA (Multimodal Interleaved Reasoning with Adaptive Visual Tokens), a novel framework that augments the conventional text vocabulary with continuous visual latent tokens distilled from vision foundation models. Our generated sequences seamlessly alternate between text and visual tokens, forming an interpretable interleaved reasoning chain. We further introduce Interleaved Reasoning Policy Optimization (IRPO), which incentivizes the model to actively select task-relevant visual tokens while preserving perceptual fidelity. Extensive evaluations demonstrate that MIRA significantly improves spatial and functional reasoning capabilities across diverse vision-centric benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.