Perception, Not Reasoning, Limits Video Spatial Understanding
Abstract
Current vision-language models (VLMs) struggle with spatial reasoning across time, and aggregate benchmark scores obscure whether failures arise from perception or reasoning. We introduce SPLIT (Separating Perception from LLM Inference with Tools), a training-free harness for video spatial reasoning that builds a metric 3-D reconstruction of a video and provides a base VLM with tools to query it. SPLIT improves accuracy on four video benchmarks (VSIBench, VSTIBench, ReVSI, and DSI-Bench) and the multi-image MMSI-Bench by up to 20.6% over the base VLM answering without tools, with gains extending to open-weight base VLMs. SPLIT helps more when a question's objects are not co-visible in any frame: on VSIBench relation questions, its improvement over the base VLM alone nearly doubles when no frame shows all the objects a question names. Accurate perception closes most of the gap to a perfect score: when the tools return ground truth (SPLIT + GT), that gap shrinks 3.6-fold on subsets of VSIBench, VSTIBench, and ReVSI, and accuracy reaches an average of 92.9%; two open-weight base VLMs nearly match it, with Gemma-4-31B averaging 90.4% and Qwen3.6-27B 88.7%. We will release the harness and code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.