ActionCache: Action-Aware Feature Caching for Interactive Video World Models
Abstract
Interactive video world models require both low-latency generation and prompt responses to user control. Existing training-free caches reuse along different axes while deciding validity from signals read inside the current model computation after the forward pass has begun. This internal gating has two limitations: (1) feature redundancy along the denoising trajectory collapses once models are distilled to three or four steps, and (2) the exogenous user action that invalidates the cache arrives before the forward pass, yet is folded into a model-side quantity and so reaches the decision only after that pass has begun. We present ActionCache, which reads the action before each forward pass and uses it as the governing signal for cache reuse: it recomputes when the control changes or the cached update grows too old, and reuses it otherwise. Because the criterion comes from the control stream, reuse survives few-step distillation. ActionCache operates directly on pretrained generators and integrates into existing interactive video world models. On three such models, spanning long diffusion sampling and distilled short-step streaming generation, ActionCache attains up to speedup while the action-response-latency increase stays within frame on every backbone. These results point to a caching principle for interactive generation: the action stream as an ahead-of-time prior over cache validity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.