acceptodds
Under review as a conference paper at ICLR 2027

FrontierSpatial: Holistic Evaluation of Spatial Reasoning in Large Multimodal Models

Abstract

Spatial reasoning is foundational to intelligent systems, yet its evaluation in multimodal large language models (MLLMs) remains fragmented across narrow benchmarks that test isolated skills. We introduce ***FrontierSpatial***, a unified benchmark and evaluation framework for large-scale spatial reasoning. We consolidate 23 existing datasets into **7,838 questions paired with 11,825 images**, organized into 7 categories, 19 sub-categories, and 65 fine-grained tasks spanning physical, geometric, knowledge-grounded, relational/causal, temporal, quantitative, and perceptual reasoning. Beyond the benchmark, we provide (1) an automated curation pipeline that distills over 60,000 raw instances into a balanced dataset and generalizes to other domains, (2) an automated evaluation harness for reproducible assessment of 16 state-of-the-art MLLMs, and (3) an automated error-analysis harness for fine-grained diagnosis across the task hierarchy. Our evaluation reveals large and uneven performance gaps across spatial skills, exposing systematic weaknesses that remain hidden in isolated benchmarks. ***FrontierSpatial*** provides a rigorous foundation for systematic, fine-grained evaluation and deeper understanding of multimodal spatial reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.