GLIMPSE: GLOBAL–LOCAL INDEXING WITH MULTI- PRECISION SPECTRAL ENCODING FOR EXPLORABLE VIDEO WORLD MODELS
Abstract
A camera-controlled video world model must keep a scene consistent when the camera returns to it. Memories that render stored features into the target frame resolve estimated depth and pose into pixel positions, where geometric error becomes ghosting or duplicated structure. Positional transport avoids that resampling by relocating a historical patch’s address rather than its content, yet still encodes the projected coordinate at full precision in every rotary band, whatever the quality of the geometry behind it. An erroneous address retains full positional amplitude and can compete with correctly aligned evidence. We present GLIMPSE, a geometry-addressed memory interface that adapts positional precision and spectral attenuation to geometric reliability. It organizes historical evidence into a geometry-addressed local memory for view-specific appearance reuse and a recurrent global memory for scene-level context, with footprint-guided retrieval under a fixed local read budget. Under matched retrieval and read budgets, GLIMPSE improves geometric correspondence and photometric consistency over deterministic positional transport.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.