Beyond Geometric Alignment: Making 3D Latents Matter for Spatial Reasoning
Abstract
Humans perform spatial reasoning by constructing internal spatial representations that support mental rotation, viewpoint transformation, and relational inference. Inspired by this ability, recent multimodal large language models introduce 3D latent reasoning, in which continuous latent tokens are expected to simulate spatial imagination steps. However, it remains unclear whether these latent tokens actually encode meaningful 3D geometry and contribute to spatial reasoning. In this work, we systematically analyze 3D latent reasoning and find that latent tokens have only weak causal influence on spatial reasoning and exhibit highly homogeneous representations across scenes. To reveal the underlying cause, we further identify an image-dominant shortcut during geometric alignment, whereby the alignment can be achieved largely from image features with limited reliance on 3D latents. Based on this diagnosis, we propose CALD-3D, a framework that combines contrastive geometric alignment with intra-latent diversity regularization to make latent tokens both geometrically meaningful and causally influential in spatial reasoning. Experiments on MindCube-Tiny and VSI-Bench demonstrate consistent improvements in spatial reasoning performance while substantially alleviating latent representation collapse.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.