Perception-Action Conditioned Efficient Inference for Interactive Video World Models
Abstract
Interactive video world models enable users and agents to continuously explore generated environments via action-conditioned visual generation, benefiting applications in game simulation, autonomous driving, and embodied intelligence. Nevertheless, the repeated Diffusion Transformer (DiT) denoising conditioned on the long visual history incurs substantial latency, severely hindering responsive interaction. Existing approaches decrease the latency through reducing redundant computations in generated context, attention and denoising, by exploring the relevance between historically generated contents and the one under the current action. However, they generally rely on limited guidance signals for compression and adopt fixed reduction schemes, overlooking perceptual redundancy and dynamiclly varying computation demands, thus suffering significant performance degradation. To address this issue, we propose Perception–Action Conditioned Efficient Inference (PACE), a training-free framework that synergizes the perceived visual evidence with action cues to dynamically allocate computation. Specifically, we firstly develop the Block-wise Context Compression method to filter historical blocks based on visual redundancy and target-view relevance. Subsequently, we employ Action-Conditioned Attention Sparsity by optimally distributing attention budgets between past and current tokens. Besides, we design the Interruption-Aware Cache Reuse strategy, by scheduling feature reuse based on visual stability while conservatively adjusting during abrupt action transitions. On HY-WorldPlay and Matrix-Game-3.0, our method achieves 2.77× and 1.74× DiT-core speedup, respectively, remarkably promoting both inference latency and peak memory overhead with improved generation fidelity in most cases, compared to the state-of-the-art approaches. We will release source code upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.