acceptodds
Under review as a conference paper at ICLR 2027

WorldKV: Efficient World Memory with World Retrieval and Compression

Abstract

Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpoint yields consistent content, remains an open problem. Full KV cache attention may preserve this consistency but breaks real-time constraints: memory footprint and attention cost grow linearly with rollout length. Sliding window inference restores throughput but discards long-term memory. We observe that, even without memory-specific training, attention over historical KV caches follows viewpoint correspondence, concentrating on chunks that overlap with the current view. Motivated by this observation, we propose WorldKV, a training-free framework for persistent world modeling with two components: World Retrieval selectively retrieves scene-relevant chunks via camera/action correspondence, inserting them back into the native attention window without re-encoding. World Compression prunes redundant tokens within each chunk via key-key similarity to an anchor frame, halving per-chunk storage to fit 2 more history under a fixed budget. On Matrix-Game-2.0 and LingBot-World-Fast, WorldKV matches or exceeds the fidelity achieved with Full KV memory, at roughly 2 the throughput.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.