acceptodds
Under review as a conference paper at ICLR 2027

DeSpatial: The devil of Agentic Spatial Reasoning Lies in 3D Itself

Abstract

Spatial agents have access to reconstruction and segmentation tools, yet inconsistent predictions and fragmented object identities can prevent their outputs from becoming reliable 3D scene information. A simple oracle experiment exposes this perception bottleneck: replacing tool-predicted perception with ground-truth 3D bounding boxes raises an agent to nearly full score across five object-centric VSI-Bench tasks, while retaining the same reasoning tools. Motivated by this finding, we believe the devil of spatial reasoning lies in 3D itself, and introduce DeSpatial(Devil of Spatial), which explicitly decomposes spatial understanding into perception and reasoning. DeSpatial has a shared 3D perception tool to reconcile independent perception outputs into a unified, metric, gravity-aligned 3D scene. This perception tool has two core components: a) a joint segmentation–depth co-refinement module that consolidates masks of the same object and removes projection artifacts caused by inconsistent segmentation and depth within a single frame; b) a cross-view instance fusion module that consolidate voxel-overlapping instances but not allowed to merge same-frame instances. This fusion module additionally fuses spatially disjoint fragments caused by insufficient multi-view observations, by checking their 3D structure and source mask appearance. After receiving the complete and unified 3D scene information, the agent chooses reasoning or computational tools to solve the spatial problem. DeSpatial achieves leading performance on VSI-Bench, MindCube-Tiny, and ViewSpatial, with additional long-video results on VSI-SUPER, revealing the bottleneck of spatial understanding lies in a precise 3D perception tool.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.