NARRA: Self-Adapting Video World Models at Test Time
Abstract
Long-horizon video world models often suffer from visual drift: as generation progresses, the appearance and geometry of a scene can gradually change, even when the camera returns to a previously observed viewpoint. We introduce NARRA, a self-adaptive framework that enables video world models to learn and remember at test time. NARRA uses an initial generation as self-supervised context to specialize the model to each generated world. Inspired by Nested Learning, adaptation occurs across interacting learning processes with different update frequencies, allowing new information to be incorporated without overwriting pretrained knowledge. The framework can further enhance video quality by leveraging channel-wise memory of earlier temporal segments and known camera trajectories that identify revisited viewpoints. We show that NARRA significantly improves video quality, with consistent gains across VBench and other evaluation metrics on the SANA-WM benchmark, and particularly strong improvements on challenging trajectories that require the model to return consistently to previously observed viewpoints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.