acceptodds
Under review as a conference paper at ICLR 2027

WorldRetain: Predictive Retention and Spatial Compression for Efficient and Consistent Video World Model Inference

Abstract

Autoregressive video world models must preserve scene identity across viewpoint revisits while limiting the cost of historical context. Sliding-window inference can discard evidence needed for later returns, and without an external memory bank, this eviction is irreversible. We propose WorldRetain, an inference-time cache policy that combines predictive KV retention with spatial compression. A lightweight predictor learns from future attention and viewpoint reuse to rank historical chunks, using only observed camera geometry and resident cache features at inference. Selected history is compressed into salient tokens and pooled summaries, maintaining a compact native cache without an external bank or changes to generator weights. On 100 RealCam-Vid camera-loop trajectories per backbone, WorldRetain achieves the best scene-identity and appearance consistency among the evaluated methods on LingBot-World-Fast and Matrix-Game 2.0. Compared with sliding-window inference, same-place scores increase from 1.77 to 2.29 and from 1.06 to 2.71, respectively, alongside 1.37–1.43× throughput speedups and lower peak GPU memory

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.