SpatialRoute-Bench: Benchmarking Long-Horizon Spatial Intelligence of Vision-Language Models in Outdoor Videos
Abstract
Vision-language models (VLMs) have made rapid progress in spatial understanding, yet existing benchmarks largely focus on local or scene-level spatial relations, leaving long-horizon spatial understanding in continuous outdoor videos underexplored. Such videos require models to integrate spatial evidence scattered across long time spans and large physical distances, as the observer's viewpoint and direction of travel change continuously along the route. To bridge this gap, we introduce SpatialRoute-Bench, a geometry-grounded benchmark for evaluating how VLMs perceive, reason over, and use spatial information along extended outdoor trajectories. SpatialRoute-Bench comprises 7,000 multiple-choice QA pairs across seven tasks, organized into a three-level hierarchy in which each level demands more extensive spatial integration than the one below: perception, reasoning, and decision. To enable scalable benchmark construction, we develop a fully automated pipeline that couples VLM-based semantic grounding with 3D reconstruction: semantic annotations determine which landmarks can be referred to, whereas reconstructed camera poses, trajectories, and scene geometry provide the spatial evidence from which answers are derived. Task-specific generators then convert this evidence into questions, retaining only those that pass geometric and semantic validity checks. Evaluations of representative VLMs reveal three critical limitations. First, there remains substantial room for improvement: even the strongest model reaches only 28.25% overall accuracy against a 25% random baseline. Second, camera-motion perception remains unreliable: even for a single motion event, average accuracy is only 26.2%. Third, spatial relations degrade over longer horizons: when the same relation must be recovered between landmarks observed minutes apart, average accuracy drops from 29.2% to 22.1%, below the random baseline. SpatialRoute-Bench exposes a clear gap between general-purpose video understanding and grounded spatial understanding of long outdoor traversals, and provides a diagnostic testbed for closing it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.