acceptodds
Under review as a conference paper at ICLR 2027

Learning Persistent World Representations for Spatial Reasoning from Partial Observations

Abstract

Spatial reasoning over videos and multiple views requires integrating spatial evidence distributed across partial observations. However, existing approaches mainly aggregate visual evidence without explicitly constructing a representation of the observed world. We argue that spatial reasoning should be mediated by a query-independent persistent world representation, causally constructed from partial observations before downstream reasoning. Following this formulation, we propose the *Internal Spatial World Encoder* (ISWE), which augments a vision-language model with a causal world-construction process that enables world-informed visual enrichment. ISWE preserves observation-specific evidence in local states and causally integrates it into persistent World representations through asymmetric Local-to-World updates. The constructed world is maintained through complementary temporal and persistent memory carriers, and is accessed to enrich visual tokens by combining local and cross-observation information before query-dependent answer generation. We further propose *Scope-Aligned Spatial Internalization*, which aligns spatial supervision with the validity scopes of local, world, and memory representations. ISWE-8B achieves a macro-average score of **69.6** on VSI-Bench. It further demonstrates strong generalization on ViewSpatial-Bench and MMSI-Bench, achieving **58.8** and **38.3**, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.