Reconstruction as Spatial Evaluation: Diagnosing Cross-View Spatial Intelligence in Large Vision-Language Models
Abstract
Spatial intelligence fundamentally requires a model to combine local visual observations into a geometrically consistent representation of the physical world. Many static benchmarks assess isolated spatial judgments, while interactive embodied evaluations combine perception with exploration, planning, and control. To bridge this gap, we introduce RECONSPATIAL, a benchmark centered on multi-view 3D scene reconstruction across indoor, tabletop, and outdoor environments. Given real multi-view images and target-object queries, large vision- language models (LVLMs) jointly predict oriented 3D bounding boxes in a shared, camera-anchored metric coordinate frame. Reconstructions are evaluated along two complementary axes: Structure for relative layout and Metric for physical scale and camera-anchored placement. Nine atomic spatial QA tasks, grouped into Geometric Perception and Spatial Relations, probe spatial skills involved in joint reconstruction. Offline navigation extends evaluation to goal-directed planning and provides a setting to examine reconstructed maps as intermediate reasoning representations. Evaluation of 18 frontier LVLMs reveals a consistent gap between relative scene structure and metric reconstruction; even GPT-6-Astra scores 81.4% on Structure but only 58.7% on Metric. Diagnostic analyses link atomic spatial skills to reconstruction quality, while showing that relational consistency alone does not ensure metric accuracy. Explicit 3D maps provide an intermediate representation that separates scene construction from path generation, revealing how geometric errors limit navigation success. Finally, spatial fine-tuning yields broad transfer across both spatial and general vision-language benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.