acceptodds
Under review as a conference paper at ICLR 2027

GS-MLLM: Cognitively Inspired Triadic Perception for Multiview Spatial Reasoning

Abstract

Recent progress in spatial MLLMs increasingly leverages geometric priors from 3D vision models. However, exposing reconstructed geometry to an MLLM does not resolve a fundamental problem: how should geometric evidence be organized for reasoning? Geometry must associate repeated observations across views without erasing view-specific details, while a limited reasoning capacity must preserve both question-relevant evidence and global scene structure. We introduce GS-MLLM, a framework for multimodal spatial reasoning that organizes reconstructed geometry into reliability-aware physical groups and combines question-conditioned local evidence with global scene context. These groups consolidate geometrically consistent observations across views while retaining their source-view evidence. A question-conditioned semantic memory retrieves task-relevant local detail, while a question-independent adaptive spatial prefix summarizes the global scene layout. Position-aware, frame-aligned cross-attention integrates the retrieved geometry into Pixel-token states while preserving the MLLM's pretrained visual-token structure. With Qwen3-VL-4B, GS-MLLM achieves 54.7 on VSI-Bench, surpassing the previous best result by 3.1 points, and delivers a further 2.6-point average gain over GeoThinker across four additional spatial reasoning benchmarks. These results suggest that the key to exploiting 3D priors in MLLMs lies not only in providing geometry, but in structuring it around the physical scene and exposing it at the appropriate level of detail.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.