SpatialAct: Probing the Gap Between Static Spatial Reasoning and Closed-Loop Action of VLM agents in 3D Scenes
Abstract
Humans exhibit remarkable spatial intelligence, enabling them to perceive spatial structure, reason about their surroundings, and adapt effectively as environments change. As frontier models such as GPT-6 Astra continue to advance, an important question arises as to how far current vision-language models (VLMs) have progressed from classical spatial understanding toward more complex interactive spatial intelligence. To study this transition, we introduce , a simulator-grounded benchmark for evaluating VLMs from static spatial reasoning to closed-loop action in 3D scenes. SpatialAct adopts Multi-turn Interactive Refinement as its primary task and connects foundational spatial abilities and single-step correction within a hierarchical evaluation framework. Experiments across leading proprietary and open-source VLMs reveal a substantial performance gap as tasks progress from isolated spatial judgments to complex multi-turn interaction. While current VLMs perform relatively well on static spatial tasks, their performance drops markedly in closed-loop refinement, where even GPT-6 Astra achieves a Repair Rate of 0.681 and a Scene Success Rate of 0.332, compared with 0.911 and 0.763 for human participants. We further investigate the sources of this cross-task capability gap by decomposing the refinement loop into spatial reasoning, corrective action, and spatial state tracking, revealing distinct limitations that emerge during interaction. Finally, we explore external support for bridging this gap and find that feedback and memory can partially compensate for possible limitations in current VLMs' internal spatial representations, particularly in anticipating action consequences and updating spatial states over time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.