SpatialAnything: Bidirectional Attention and Multi-Teacher Distillation for Embodied Spatial Reasoning
Abstract
Large vision-language models (LVLMs) recently demonstrate promising capabilities in embodied spatial reasoning, spanning spatial understanding, object localization, navigation, and long-horizon planning. These tasks require models to process both language and visual grounding primitives such as bounding boxes, keypoints, and trajectories. However, existing methods typically generate coordinates one at a time through autoregressive decoding, so coordinates within the same visual grounding primitive depend on different contexts and generation time grows with primitive length. Meanwhile, different spatial tasks build on shared localization and spatial relation understanding, yet differ in output formats and optimization objectives. We therefore present SpatialAnything, a natively multimodal model for embodied spatial reasoning built around Primitive Bidirectional Attention (PBA). PBA enables bidirectional attention among coordinate representations within a primitive through bidirectional masks in full-attention layers and query rewriting with write compensation in gated linear-attention layers. SpatialAnything further adopts adaptive parallel decoding, which treats each primitive as a whole and predicts all of its coordinates at once, removing the wait inherent in coordinate-by-coordinate generation. The predictions are then written into the context to support subsequent generation. Building on this, SpatialAnything uses multi-task joint training through a two-stage strategy based on Spatial Multi-Teacher On-Policy Distillation (Spatial MOPD), which first establishes a foundation for localization prediction with a unified primitive teacher and then lets task-specific teachers guide the student on its own generations, giving a single model complementary spatial capabilities. Across a range of embodied spatial reasoning benchmarks, SpatialAnything achieves the highest scores on multiple metrics, exceeding GPT 6 Astra by 5.2 points on average over ten spatial understanding and reasoning benchmarks, and delivers a 2.9× inference speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.