SpatialWeave: Co-Evolving Tools, Workflows, and Skills for Spatial Reasoning
Abstract
Tool-augmented vision-language models can perceive complex spatial environments, but they struggle to calculate the exact answers that questions require. Existing methods either recalculate everything from scratch for each question and repeat the same mistakes, or they use fixed steps that fail on new problems. We introduce SpatialWeave, which improves spatial reasoning without modifying the underlying language model. Instead, it automatically optimizes the external tools, workflows, and usage rules the model relies on. During evolution, a temporary procedure scaffold guides the model to explore various tool combinations. Successful tool sequences are saved as reusable workflows, and common mistakes are corrected by adding specific skills. After evolution, this scaffold is completely removed. Across seven spatial reasoning benchmarks, SpatialWeave consistently improves performance on three different models. Notably, it raises Qwen3.8-27B to 80.73%, beating the base GPT-6 Astra model at 80.28%. It also improves Gemma-4-31B to 73.58%, outperforming the specialized SpatialClaw agent on the same model. Furthermore, removing the scaffold during inference makes the process much more efficient. It reduces token generation by up to 34.9% and tool calls by up to 46.0% without losing accuracy. Ultimately, SpatialWeave shows that automatically building a lightweight set of external tools and rules is a practical way to achieve strong spatial reasoning while keeping inference fast and simple.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.