acceptodds
Under review as a conference paper at ICLR 2027

Learning a Spatial Memory Codec for Video World Model

Abstract

Generative video world models are emerging as a foundation for interactive agents, simulators, and creative tools that synthesize visual trajectories through spatial environments. Scaling such models to long trajectories requires more than local temporal coherence: when a place is revisited, the model must recover visual information from much earlier observations. Since keeping all observations in the model's context is not scalable, the model needs a long-term memory that preserves useful visual information while keeping storage compact and retrieval efficient. Existing approaches summarize history into a fixed-size state, retain a growing archive, or rely on explicit 3D reconstruction. These choices can lose important visual information, incur high storage and retrieval costs, or require hand-crafted update rules. We introduce a learned spatial memory codec that stores history in compact learned entries. These entries are retrieved by a reader through a hierarchical, pose-indexed mechanism that efficiently selects those most relevant to each generated view. Specifically, each video chunk is compressed into a small set of entries, each paired with a compact index representation that we call a landmark. Each landmark summarizes its corresponding entry and records where and when it was observed. To form the generation context, the reader first selects a fixed-size set of the most relevant landmarks and then uses only their corresponding entries. We show that our memory codec balances compact storage, efficient retrieval, and consistent generation across challenging revisits, including viewpoint changes and full U-turns. We further demonstrate retrieval from nearly two hours of observations, far beyond the training horizon.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.