Learning Persistent World Representations for Spatial Reasoning from Partial Observations
Abstract
Spatial reasoning over videos and multiple views requires integrating spatial evidence distributed across partial observations. However, existing approaches mainly aggregate visual evidence without explicitly constructing a representation of the observed world. We argue that spatial reasoning should be mediated by a query-independent persistent world representation, causally constructed from partial observations before downstream reasoning. Following this formulation, we propose the *Internal Spatial World Encoder* (ISWE), which augments a vision-language model with a causal world-construction process that enables world-informed visual enrichment. ISWE preserves observation-specific evidence in local states and causally integrates it into persistent World representations through asymmetric Local-to-World updates. The constructed world is maintained through complementary temporal and persistent memory carriers, and is accessed to enrich visual tokens by combining local and cross-observation information before query-dependent answer generation. We further propose *Scope-Aligned Spatial Internalization*, which aligns spatial supervision with the validity scopes of local, world, and memory representations. ISWE-8B achieves a macro-average score of **69.6** on VSI-Bench. It further demonstrates strong generalization on ViewSpatial-Bench and MMSI-Bench, achieving **58.8** and **38.3**, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.