MindTraverse: Benchmarking Cross-Scale Spatial Composition
Abstract
Real-world spatial tasks often exhibit **cross-scale structure**: models must organize and propagate spatial context across continuous sequences of local spatial judgments and global environmental reasoning. Yet such cross-scale spatial composition has been largely overlooked by existing benchmarks, which either focus on single-step reasoning or fail to capture the cross-scale nature of real-world space in multi-step tasks. To investigate this underexplored problem, we construct **MindTraverse**. The benchmark recursively composes atomic task interfaces to create multi-step cross-scale reasoning problems validated through both geometric and visual checks. In MindTraverse, models must continuously propagate and reason over spatial context between fine-grained, object-level workspace reasoning and coarse-grained, environment-level transit reasoning. The resulting benchmark contains **6,167 multi-view questions**, covering **35 composition templates**, **six scale sequences**, and **210 simulator scenes**. Across **20 model configurations**, even the best-performing model achieves only **50.92% overall accuracy**. More importantly, endpoint accuracy does not decrease monotonically as the intended compositional dependency becomes deeper. To investigate this counterintuitive behavior, we conduct further controlled experiments and reveal three systematic deficiencies in current spatial reasoning models. First, **Structural Binding Failure**: models recover task-relevant semantic entities but struggle to organize them into coherent cross-scale structure. Second, **State–Evidence Conflict**: even when explicit spatial state is available, visual evidence is not integrated reliably and can interfere with downstream reasoning. Third, **Trajectory-Role Shortcut**: on some high-accuracy tasks, models bypass the intended cross-scale dependency and collapse onto simple trajectory roles. These results show that genuine cross-scale spatial reasoning requires not only correct final answers, but also correct spatial representation, reliable evidence integration, and faithful downstream use of that representation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.