Persistent4D: A Training-Free Streaming World Model with persistent 4D memory
Abstract
A visual world model must let an agent apply an action and predict spatio-temporally consistent future observations, which requires both precise camera-controllable generation and a spatial memory that persists across steps. Existing approaches each solve only half of the problem. Controllable-video and interactive world models typically take a strongly pre-trained video generator and further train or fine-tune an action-/camera-control signal on top of it, which demands substantial compute and large paired video–action datasets, and still lacks a persistent geometric memory—so long rollouts drift and revisited viewpoints are not reproduced. Feed-forward 4D reconstruction, in contrast, is training-free but can only re-project already-observed content and cannot hallucinate the disocclusions exposed by a new viewpoint. We present Persistent 4D Memory, a fully training-free streaming world model—we train, fine-tune, and distill nothing, not even a camera-control adapter—that couples an off-the-shelf feed-forward reconstructor with an off-the-shelf video-diffusion inpainter through a cross-round, per-frame 4D point-cloud memory. Each round renders the memory from the target camera (camera control by construction), lets the off-the-shelf generator fill the shrinking disocclusion holes, then re-reconstructs the completed clip under the known camera and fuses the newly revealed content back into memory. This yields three properties: (i) near-zero camera-following error without any camera-conditioning training; (ii) cross-round fusion makes the rollout cycle-consistent—turning away and back reproduces the original view, with hole area dropping from about 58% for a memory-less baseline to under 12%; and (iii) because geometry offloads the vast majority of content and the generator fills only small holes, a single off-the-shelf (community-accelerated) generator suffices for real time—up to ∼20 FPS with a causal streaming back-end—while we train nothing ourselves. We further propose cycle-consistency as an evaluation protocol for world models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.