Recognition Gains, Compositional Gaps: Rethinking Spatial Reasoning in Vision-Language Models
Abstract
Spatial reasoning underpins real-world applications of vision-language models, including robotics, autonomous driving, and navigation. Yet progress in spatial recognition leaves unresolved whether models can reliably combine and reuse the information they recognize. To examine this relationship, we introduce SpaReCo-Bench (Spatial Recognition and Composition Benchmark), which organizes 15 tasks across four families: state, identity, relation, and composition. Centered on object orientation and spatial transformations, the benchmark grounds basic judgments and compositional problems in shared visual contexts, making it possible to assess recognition and the use of recognized information together. Our evaluations reveal a persistent gap between the two: models can correctly identify the spatial states and relations needed for a task, yet fail when these judgments must be combined. This disconnect persists after spatial training, even as recognition and the application of explicit transformations improve, showing that gains in these capabilities do not by themselves resolve difficulties in attribute composition and transformation analogy. SpaReCo-Bench provides a controlled foundation for distinguishing recognition gains from progress in compositional reasoning and for diagnosing where spatial information ceases to support reliable task performance. The benchmark, dataset, and code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.