Learning to Remember and Evolve: Object State Memory for Video World Models
Abstract
Camera- and text-controlled video world models require persistent scene information to connect generated changes across viewpoints and beyond recent video context. We propose a framework that connects online feed-forward Gaussian reconstruction, persistent state maintenance, and spatiotemporal conditioning in a recurrent generation pipeline. StateSplat provides renderable geometry and object-level visual evidence, which learned memory integrates alongside explicit state records. Observed and accepted generated records remain distinct from future predictions. Time- and camera-dependent readout supplies memory to the generator through temporal tokens and spatial feature injection, while admitted generated frames provide evidence for subsequent updates. Video generation and subsequent-observation reconstruction objectives jointly train feature adaptation, memory updating, and readout. Experiments on CARLA and DAVIS show improved object reappearance and reference-based continuation quality, with ablations supporting the combined memory configuration and spatial conditioning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.