FrontierSpatial: Holistic Evaluation of Spatial Reasoning in Large Multimodal Models
Abstract
Spatial reasoning is foundational to intelligent systems, yet its evaluation in multimodal large language models (MLLMs) remains fragmented across narrow benchmarks that test isolated skills. We introduce ***FrontierSpatial***, a unified benchmark and evaluation framework for large-scale spatial reasoning. We consolidate 23 existing datasets into **7,838 questions paired with 11,825 images**, organized into 7 categories, 19 sub-categories, and 65 fine-grained tasks spanning physical, geometric, knowledge-grounded, relational/causal, temporal, quantitative, and perceptual reasoning. Beyond the benchmark, we provide (1) an automated curation pipeline that distills over 60,000 raw instances into a balanced dataset and generalizes to other domains, (2) an automated evaluation harness for reproducible assessment of 16 state-of-the-art MLLMs, and (3) an automated error-analysis harness for fine-grained diagnosis across the task hierarchy. Our evaluation reveals large and uneven performance gaps across spatial skills, exposing systematic weaknesses that remain hidden in isolated benchmarks. ***FrontierSpatial*** provides a rigorous foundation for systematic, fine-grained evaluation and deeper understanding of multimodal spatial reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.