MAX-Cache: Motion-Aligned Activation Reuse with Mismatch-Guided Selective Recomputation for Efficient Autoregressive Video Diffusion
Abstract
Interactive driving simulation requires video world models to generate future scenes continuously and with low latency. Existing acceleration methods reduce denoising steps, sparsify attention, or optimize key-value caches. We exploit a complementary source of redundancy: neighboring video chunks share scene content and have similar intermediate activations. Reusing these activations at unchanged token locations, however, is sensitive to ego motion, independently moving objects, parallax, and occlusion. We proposeMAX-Cache, an inference-acceleration framework with lightweight post-hoc calibration for motion-aligned activation reuse across chunks. MAX-Cache evaluates a safety gate before either reuse mechanism. When the gate passes, it directly reuses the entire motion-aligned output at , bypasses the prefix through , and repairs that feature with a lightweight residual corrector calibrated post hoc from training-split teacher pairs while the pretrained generator remains frozen. Otherwise, the executed blocks use a partial token mask: low-mismatch tokens reuse warped cached outputs, while high-mismatch tokens retain fresh block updates. Conservative reuse criteria and baseline K/V-state maintenance limit error accumulation. On the 21-frame single-view nuScenes setting, the resulting configuration increases end-to-end throughput from 3.90 to 5.48 FPS (1.41). Across both 21- and 211-frame settings, VBench, FID, and FVD remain close to the baseline, while the measured 211-frame end-to-end speedup reaches 1.30.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.