ConsistKV: Long-Term KV Memory for Consistent Interactive World Models
Abstract
Interactive world models can generate controllable visual environments in real time, yet they often fail to preserve previously observed scenes over long interactions. We identify this failure as revisit collapse: once historical observations leave the bounded KV context, the model loses direct access to their visual states and may reconstruct different content when revisiting the same location. We propose ConsistKV, a lightweight external KV memory for frozen interactive world models that maintains compressed distant history without extending the active attention window. For retrieval, our Action–Content Encoder (ACE) jointly models visual content and action-derived relative geometry, combining appearance correspondence with complementary geometric localization. We further introduce a training-free Value-Absorbing KV Compression strategy that reduces historical KV storage while preserving information from discarded positions through shared self-Gram-guided selection and value absorption. Experiments on Matrix-Game 2.0 and ABot-World show improved revisit consistency over sliding-window and fixed-budget memory baselines, with clearer gains over longer interaction horizons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.