ActCache: Action-Guided Cross-Chunk KV Cache Reuse for Interactive World Models
Abstract
Interactive world models generate continuous visual environments chunk by chunk while maintaining historical KV caches for temporal coherence. Although few-step distillation reduces denoising iterations, dense computation within each new chunk remains a major barrier to real-time interaction, while conventional step-wise caching offers limited reuse opportunities. We observe that adjacent chunks preserve substantial visual content and that explicit controls constrain its motion, making historical KV representations a structured source of reusable computation. Based on this insight, we introduce **ActCache**, an action-guided framework for fine-grained cross-chunk KV reuse. ActCache converts motion-related controls into coarse displacement priors and restricts correspondence search to layer-adaptive neighborhoods. Since these priors cannot fully capture parallax, object motion, occlusion, or newly revealed content, ActCache uses the initial clean estimate from the first denoising step as a self-generated probe to verify token-level correspondences. Verified tokens reuse historical self-attention representations, while unmatched regions are recomputed. A shift-equivariant calibration mechanism applies relative 3D-RoPE rotations to align reused keys with their current coordinates. Without additional training, ActCache delivers up to speedup across three interactive world models while preserving world-memory consistency. Combined with complementary acceleration strategies, it further achieves up to speedup.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.