acceptodds
Under review as a conference paper at ICLR 2027

SpatialAct: Probing the Gap Between Static Spatial Reasoning and Closed-Loop Action of VLM agents in 3D Scenes

Abstract

Humans exhibit remarkable spatial intelligence, enabling them to perceive spatial structure, reason about their surroundings, and adapt effectively as environments change. As frontier models such as GPT-6 Astra continue to advance, an important question arises as to how far current vision-language models (VLMs) have progressed from classical spatial understanding toward more complex interactive spatial intelligence. To study this transition, we introduce , a simulator-grounded benchmark for evaluating VLMs from static spatial reasoning to closed-loop action in 3D scenes. SpatialAct adopts Multi-turn Interactive Refinement as its primary task and connects foundational spatial abilities and single-step correction within a hierarchical evaluation framework. Experiments across leading proprietary and open-source VLMs reveal a substantial performance gap as tasks progress from isolated spatial judgments to complex multi-turn interaction. While current VLMs perform relatively well on static spatial tasks, their performance drops markedly in closed-loop refinement, where even GPT-6 Astra achieves a Repair Rate of 0.681 and a Scene Success Rate of 0.332, compared with 0.911 and 0.763 for human participants. We further investigate the sources of this cross-task capability gap by decomposing the refinement loop into spatial reasoning, corrective action, and spatial state tracking, revealing distinct limitations that emerge during interaction. Finally, we explore external support for bridging this gap and find that feedback and memory can partially compensate for possible limitations in current VLMs' internal spatial representations, particularly in anticipating action consequences and updating spatial states over time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.