DM3SR: Benchmarking Deep Multi-View 3D Spatial Reasoning
Abstract
Vision–language models show strong VQA performance but remain weak at multi-view 3D spatial reasoning. Many existing benchmarks, particularly those built from video-derived frames, provide limited coverage of viewpoint diversity and vertical scene structure, leaving model behavior under broader 3D configurations insufficiently characterized. We introduce DM3SR (Deep Multi-View 3D Spatial Reasoning), a geometry-controlled benchmark spanning three complementary scene families: Contextual, Controlled-Syn and Controlled-Real. DM3SR defines twelve subtasks covering object counting, absolute and relative distance, and lateral and vertical positional reasoning. Evaluating 32 VLMs reveals substantial gaps from human performance, task-dependent benefits from scaling, and inconsistent transfer of spatial adaptation across scenes and axes. Controlled analyses indicate preferences for XY-dominated geometry and continuous viewpoints, alongside difficulty integrating local observations along dimensions with greater scene extent. DM3SR provides a testbed for evaluating cross-view integration and diagnosing how scene geometry and viewing conditions shape spatial reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.