TokenWeave: Consolidating Temporal KV History into Persistent Scene Memory for Video World Models
Abstract
Streaming autoregressive video world models naturally accumulate history as a temporal sequence of key–value (KV) states. However, long interactions repeatedly observe the same scene content across distant times, making such temporally organized history increasingly redundant and eventually forcing old observations to be discarded. We introduce TokenWeave, a training-free framework that reorganizes temporal KV history into persistent scene memory by associating and consolidating repeated patch-level observations. Rather than managing history as a collection of separate observations, we reorganize repeated observations around recurring scene content. It operates entirely on the frozen model’s internal KV states, grouping observations that likely correspond to the same scene content and integrating them into persistent memory entries. These shared entries are updated with position- and multiplicity-aware fusion, while genuinely novel content is retained as new memory. Memory therefore grows mainly with new or unmatched content rather than with time. % The resulting memory therefore scales with scene novelty rather than rollout duration. Experiments across multiple domains and rollout lengths demonstrate substantial improvements in long-horizon revisit consistency without memory-specific training. Additional videos are available at https://tokenweave-iclr.github.io/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.