SpaceMCP: Vision Language Model Spatial Reasoning via Model Context Protocol
Abstract
Despite impressive advances in large language models (LLM), vision–language models (VLM) still exhibit limited spatial reasoning capabilities. Recent progress in feed-forward 3D reconstruction has enabled precise, real-time geometric understanding from visual inputs. In this work, we propose SpaceMCP, a framework that leverages these advances to introduce a Model Context Protocol (MCP) style paradigm for spatial reasoning. Our approach augments a VLM with structured function calls, allowing it to query a scene graph for spatial information such as distances, object positioning, and camera properties. The scene graph is constructed in real time using feed-forward 3D reconstruction to recover metric geometry, combined with open-vocabulary segmentation to identify and localize relevant objects within the scene. Paired with VLM's strong reasoning capabilities, the VLM is able to combine this primitive information to solve more complex spatial tasks. Across nine VLMs and without any training, this approach improves scores by an average of 16 points on VSI-Bench, 6 on MMSI-Bench and 14 on VSTI-Bench, and with matched backbones it outperforms the strongest training-free tool-using agent on VSI-Bench. Gains are largest on metric and camera-geometry questions, while questions that hinge on viewpoint conventions, temporal order or fine measurements remain hard. Given the rapid advances in feed-forward 3D computer vision, we propose this paradigm as an augmentation to VLM spatial reasoning rather than a replacement: the vision models contribute precise metric answers and the VLM context and nuance, and improving native VLM spatial reasoning remains necessary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.