Compositional Spatial Reasoning for Multimodal Large Language Models
Abstract
Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), requiring models to infer complex semantic, geometric, and temporal relationships from visual observations. Recent approaches enhance spatial reasoning by enriching visual representations with additional spatial priors from 3D reconstruction or depth estimation models. However, we find that this uniform representation alone does not guarantee better reasoning, as different queries require different compositions of spatial evidence. In this work, we identify two key challenges for effective spatial reasoning: comprehensive representation and selective evidence composition. We introduce CoSpace, a compositional spatial reasoning framework that follows the principle of representing comprehensively and reasoning selectively. CoSpace organizes heterogeneous visual information into structured semantic, geometric, and temporal representations, and performs query-adaptive evidence composition through hierarchical selection, allocation, and retrieval to construct query-relevant reasoning states. Extensive experiments across diverse spatial reasoning benchmarks demonstrate that CoSpace improves spatial reasoning performance, effectively exploits heterogeneous visual evidence, and enables query-adaptive composition of interpretable spatial representations. The code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.