Spatial-JARVIS: A Meta-Level Spatial Cognitive Regulator for Interactive Spatial Agents
Abstract
Multimodal large language models (MLLMs) have made significant progress in passive spatial reasoning, yet they still face substantial challenges in long-horizon active spatial interaction tasks. During such interactions, inaccurate beliefs about the environment and task progress can lead to erroneous decisions and persist across subsequent steps, ultimately resulting in task failure. We attribute these failures to insufficient spatial metacognition, i.e., the ability to monitor and regulate evolving cognitive states and decision-making processes. To address this limitation, we propose Spatial-JARVIS, a Meta-Level Regulator decoupled from the policy model. At each interaction step, Spatial-JARVIS monitors the agent's inferred cognitive state and evaluates the proposed decision before execution, generating corrective feedback to guide the policy model in revising its mistaken beliefs and adjusting its subsequent behavior. To learn this capability, we construct the MetaSpatial-35K dataset and adopt a two-stage training strategy consisting of SFT and MC-GRPO. We further introduce MetaSpatial-Bench, a static benchmark designed to directly evaluate models' monitoring and correction capabilities in spatial interaction. Experimental results show that Spatial-JARVIS outperforms the strongest compared MLLMs by 14.6% and 13.3% in monitoring and control accuracy on MetaSpatial-Bench, respectively. Moreover, Spatial-JARVIS improves the task success rates of six different policy models by an average of 7.7% on SpatialWorld, while reducing ineffective exploration and erroneous operations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.