acceptodds
Under review as a conference paper at ICLR 2027

SpatialCrew: Multi-Agent Framework for Reliable Tool-Augmented Spatial Reasoning

Abstract

Tool augmentation can improve the spatial reasoning of vision-language models (VLMs), but erroneous intermediate outputs can propagate through otherwise valid reasoning chains. Our empirical analysis identifies this error cascade as a major source of failure and shows that verification becomes less effective as reasoning spans more tools and domains. We propose SpatialCrew, a multi-agent framework that combines domain-specific Write–Execute–Verify–Adjust (WEVA) loops with coordinated reasoning across perception, geometry, and temporal specialists. This design keeps local verification focused while checking consistency between domains. On 20 benchmarks with Gemma4-31B, SpatialCrew reduces the error cascade rate (ECR) from the five tool-augmented baselines' mean of 78.0% to 23.4%, a 54.6-percentage-point reduction. Its overall benchmark macro-average is 62.6, exceeding the mean of all six baselines by 9.6 points and the strongest baseline by 2.7 points. Across all five evaluated backbones, SpatialCrew achieves a mean score of 62.9, an average gain of 9.1 points over the corresponding no-tool models, and a mean ECR of 24.7%. These results highlight intermediate-output quality control as a practical route to reliable tool-augmented spatial reasoning. Code: https://anonymous.4open.science/r/SpatialCrew-B650.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.