WorldKV: Efficient World Memory with World Retrieval and Compression
Abstract
Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpoint yields consistent content, remains an open problem. Full KV cache attention may preserve this consistency but breaks real-time constraints: memory footprint and attention cost grow linearly with rollout length. Sliding window inference restores throughput but discards long-term memory. We observe that, even without memory-specific training, attention over historical KV caches follows viewpoint correspondence, concentrating on chunks that overlap with the current view. Motivated by this observation, we propose WorldKV, a training-free framework for persistent world modeling with two components: World Retrieval selectively retrieves scene-relevant chunks via camera/action correspondence, inserting them back into the native attention window without re-encoding. World Compression prunes redundant tokens within each chunk via key-key similarity to an anchor frame, halving per-chunk storage to fit 2 more history under a fixed budget. On Matrix-Game-2.0 and LingBot-World-Fast, WorldKV matches or exceeds the fidelity achieved with Full KV memory, at roughly 2 the throughput.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.