HippoSpatial: Persistent Spatial Scene Graphs for Multi-View Spatial Reasoning
Abstract
Multi-view spatial reasoning requires integrating partial observations from multiple viewpoints into a coherent global understanding of object positions, relations, and perspectives, a task on which Vision-Language Models (VLMs) consistently perform near random chance. Existing training-free methods either estimate object positions without metric grounding or re-derive geometry from scratch at every reasoning step, leaving no persistent spatial representation that spans all available views. Inspired by two principles of hippocampal spatial memory: pairwise relational encodings between objects enabling inference about pairs never directly co-observed, and the strict separation between allocentric long-term memory and egocentric working memory combined only through an explicit transform at retrieval time, we propose HippoSpatial, a training-free framework for multi-view spatial reasoning. Given a question and its candidate answers, HippoSpatial uses perception tools to segment relevant objects, recover their metric 3D positions in a shared world frame, and encode pairwise spatial relations in a Spatial Scene Graph, while separately maintaining camera geometry for each input view. At query time, a VLM performs visual grounding to resolve object references to graph nodes, then dispatches deterministic geometric tools that combine allocentric object positions with egocentric viewpoint information, delegating all numerical computation away from LLM reasoning. Experiments on MindCube, MMSI-Bench, and SPAR-Bench demonstrate that HippoSpatial outperforms prior training-free methods in most backbone-benchmark settings, with gains of up to 7.0 points over the best prior method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.