LaST: Latent State Tracking for Persistent Object Memory in Video World Models
Abstract
Video world models must preserve object identity across interruptions in visibility. Yet the observations establishing that identity can leave the local attention window long before an object reappears. Persistent memory must retain relevant visual information and make it available to later generation. We introduce LaST, Latent State Tracking, which separates persistent object memory from local video generation. A recurrent writer maintains scene context, object geometry, and appearance across context windows. A learned reader uses geometry to retrieve appearance at the requested locations and combines decoded scene and object fields to correct a frozen generator's denoising predictions. We pretrain memory through future reconstruction, then jointly optimize state construction and reader corrections over multiple generation windows. Parameter-free rescaling regulates spatial contrast. On the full WRBench evaluation set, LaST improves all five metrics over both KV baselines at both 1.3B and 5B. A capacity study on MOSEv2 relates reconstruction quality and object coverage to the number of retained slots. These results support retaining and reusing object history beyond the generator's local attention window.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.