In3D-VLM: Faithful Integration of Explicit 3D Space for Metric Spatial Reasoning
Abstract
As spatial reasoning becomes increasingly important for vision-language models (VLMs), incorporating explicit 3D geometry has attracted growing attention. However, existing approaches either rely on computationally expensive geometry encoders or mix compressed geometry with RGB features without ensuring that the injected 3D information remains accessible after fusion. Consequently, even when metric 3D observations are available, VLMs may struggle to accurately infer quantitative properties such as depth, distance, and relative displacement. We argue that effective 3D space injection depends on two fundamental design choices: what geometric representation to inject and how to fuse it with visual features while keeping the injected geometry explicitly accessible. To this end, we introduce In3D-VLM, which directly incorporates metric 3D geometry obtained from depth sensors or metric depth estimation models. The Point-Map Tokenizer (PMT), a lightweight tokenizer pretrained through masked XYZ point-map reconstruction, converts dense observations into compact, geometrically recoverable tokens without relying on a large geometry backbone. Geometry-Preserving Fusion (GPF) then maps these geometry tokens through a fixed orthonormal basis into a dedicated subspace of the LLM embedding space and adds them to the spatially aligned RGB visual tokens. The orthonormal mapping expands the geometry tokens into the LLM space without deforming their learned geometric structure. Geometry-leak regularization further reduces semantic overlap with the reserved subspace, encouraging an explicit linear geometry readout from the fused tokens. We further introduce MD2Bench to evaluate metric depth and distance reasoning from provided 3D geometry under visual-point and text-name grounding protocols. In3D-VLM substantially outperforms existing VLMs on MDBench and achieves state-of-the-art performance on established public spatial-reasoning benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.