acceptodds
Under review as a conference paper at ICLR 2027

SpaceMCP: Vision Language Model Spatial Reasoning via Model Context Protocol

Abstract

Despite impressive advances in large language models (LLM), vision–language models (VLM) still exhibit limited spatial reasoning capabilities. Recent progress in feed-forward 3D reconstruction has enabled precise, real-time geometric understanding from visual inputs. In this work, we propose SpaceMCP, a framework that leverages these advances to introduce a Model Context Protocol (MCP) style paradigm for spatial reasoning. Our approach augments a VLM with structured function calls, allowing it to query a scene graph for spatial information such as distances, object positioning, and camera properties. The scene graph is constructed in real time using feed-forward 3D reconstruction to recover metric geometry, combined with open-vocabulary segmentation to identify and localize relevant objects within the scene. Paired with VLM's strong reasoning capabilities, the VLM is able to combine this primitive information to solve more complex spatial tasks. Across nine VLMs and without any training, this approach improves scores by an average of 16 points on VSI-Bench, 6 on MMSI-Bench and 14 on VSTI-Bench, and with matched backbones it outperforms the strongest training-free tool-using agent on VSI-Bench. Gains are largest on metric and camera-geometry questions, while questions that hinge on viewpoint conventions, temporal order or fine measurements remain hard. Given the rapid advances in feed-forward 3D computer vision, we propose this paradigm as an augmentation to VLM spatial reasoning rather than a replacement: the vision models contribute precise metric answers and the VLM context and nuance, and improving native VLM spatial reasoning remains necessary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.