GeoMem3D: Learning to Reason over Structured 3D World Memory
Abstract
Multimodal Large Language Models (MLLMs) have significantly advanced 3D spatial reasoning, yet their ability to maintain consistent spatial understanding across viewpoints remains limited. Recent methods improve spatial reasoning by extracting evidence from continuous visual observations or interactive exploration. However, this paradigm is constrained by observation-level evidence, limiting consistent spatial reasoning across views and queries. Motivated by this observation, we propose GeoMem3D, a learning framework that enables MLLMs to reason over a persistent and structured 3D world memory of all objects and their spatial relations. GeoMem3D reconstructs multi-view scenes within a shared geometric frame and organizes object entities, geometric attributes, and spatial relations into a query-independent structured 3D world memory, which provides persistent and reusable spatial knowledge for reasoning. Leveraging this memory, we explore two complementary reasoning paradigms: GeoMem-Holistic directly reasons over the complete structured memory, while GeoMem-Agentic adaptively accesses fine-grained spatial evidence from the memory during reasoning. We further introduce a two-stage learning strategy that combines supervised fine-tuning with reinforcement learning to establish effective memory utilization behaviors. Experiments on VSI-Bench-Tiny and MindCube-Tiny demonstrate consistent improvements across diverse spatial reasoning tasks, validating the effectiveness and generalizability of structured 3D world memory for spatial reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.