Understand, Imagine, then Reason: Thinking with Foresight in Unified Latent Space
Abstract
Multimodal Large Language Models (MLLMs) have made substantial progress in visual understanding, yet reasoning about spatial relationships, object interactions, and state changes remains challenging. Thinking about images primarily operates in text, thinking with images often relies on external tools for visual evidence, and existing latent visual reasoning methods largely focus on static representations. Unified multimodal models offer an opportunity to simulate scene dynamics, yet how to leverage these dynamics for visual reasoning remains underexplored. We propose **Understand-Imagine-Reason (UnI-R)**, a framework that feeds internally simulated future dynamics back into reasoning, forming an understanding, imagination, and reasoning loop within a unified latent space. We first introduce Anchor-Conditioned Dynamic Coverage to select complementary future dynamics beyond the observed static scene, while a Foresight-to-Reasoning Bridge aggregates and integrates them into the reasoning stream. During training, Consistent and Discriminative Foresight Learning promotes consistency across rollouts and discrimination across inputs. Contrastive Future-Grounded Self-Distillation further aligns predictions based on imagined futures with teacher predictions grounded in actual future observations while contrasting them with those induced by counterfactual futures. UnI-R empirically improves visual reasoning across both dynamic-centric and general scenarios, achieving an average gain of 8.76% on BabyVision and an 8.34% improvement on MME-RealWorld, demonstrating the promise of leveraging dynamics for reasoning in unified latent space.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.