LOCI: Spatial Linear Memory for Streaming World Models
Abstract
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key–value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, making individual past observations no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that maintains both representations. In half of the transformer blocks, main attention keeps a key–value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so that viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks, supplying their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, which keeps memory constant, it streams long videos and remains more faithful than full softmax under the same budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.