acceptodds
Under review as a conference paper at ICLR 2027

Thinking with Latent Geometry Tokens for Multi-Image Spatial Reasoning

Abstract

Multi-image spatial reasoning remains challenging for multimodal large language models (MLLMs). A key bottleneck lies in integrating partial observations into coherent internal spatial representations, a capability known as spatial mental modeling. Existing approaches often rely on dense spatial annotations or external tools, limiting scalability or increasing inference overhead. To reduce these dependencies, recent work explores modeling spatial representations in latent space. However, effectively encoding geometric information in latent space and incorporating it during reasoning is underexplored. We propose Thinking with Latent Geometry Tokens, a two-stage framework that grounds latent states in compact geometric representations and encourages flexible latent-mode entry. Specifically, we use camera and scene tokens from a pretrained 3D foundation model to supervise latent states, providing viewpoint-conditioned and cross-view geometry information while jointly training on text generation. We then augment Group Relative Policy Optimization (GRPO) with latent-mode entry point exploration, using successful exploratory trajectories to guide policy updates. Experiments on three multi-image benchmarks demonstrate our consistent gains over spatial reasoning baselines, with further improvements in single-image settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.