GraphMem: A Training-free Graph Retrieval and Memory Caching Framework for Consistent World Generation
Abstract
We present GraphMem, a training-free, graph-structured visual memory framework for long-horizon autoregressive video world models. Rolling KV caches enable efficient chunk-wise generation but discard earlier scene evidence once it leaves the recent window, causing inconsistency when the camera revisits a previously observed region. Full-history conditioning retains this evidence but makes the active context and inference cost grow with video length. Our key idea is to dynamically retrieve relevant history based on the current state. GraphMem organizes pixel-space history into a Vision Memory Graph. Conditioned on the current generation state, dynamic graph retrieval activates a compact set of relevant and complementary historical views. The retrieved history is encoded with timestep-aligned noise and prefilled through the frozen backbone, producing layer-wise history KV caches in its native attention space. Dynamic retrieval keeps the active context bounded, while strided prefill amortizes the additional computation. Across diverse backbones, long-horizon memory performance improves by about 2% on average and shows better revisit consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.