Complementary Future Token Pruning for World Action Models
Abstract
World action models (WAMs) jointly predict future visual states and robot actions, but processing dense future-frame tokens incurs substantial inference latency. We propose CF-Prune, a training-free framework for pruning future-frame tokens in pretrained WAMs. Our key insight is that future visual predictions support ac- tion generation and therefore need not preserve every visual detail required for high-fidelity video synthesis. CF-Prune combines instruction and action attention to identify task-relevant tokens during the initial denoising forward pass, even when future-frame and action representations remain highly noisy. An entropy- based phase identification mechanism further retains dense computation during precision phases to mitigate pruning-induced performance degradation. Exper- iments on RoboLab, RoboCasa, and real-world manipulation tasks demonstrate consistent inference speedups. At a 35% pruning ratio, CF-Prune achieves average speedups of 1.19–1.27× while improving task success rates. At a 50% pruning ratio, it achieves 1.29–1.33× average speedups on both simulation benchmarks, with overall success rates within 1.0 percentage point of the baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.