FAVP: Future-Aware Vision Token Pruning for Unified World Action Models
Abstract
Vision token pruning can accelerate inference in unified World Action Models (WAMs), yet existing approaches primarily use semantic, temporal, or action-related cues without explicitly exploiting the future-prediction knowledge learned during world modeling. To bridge this gap, we propose **FAVP**, a training-free **F**uture-**A**ware **V**ision Token **P**runing framework that reuses this capability to guide vision token selection during action inference. FAVP introduces three modules: (1) **Future-Prediction Query Augmentation** to efficiently extract future-aware visual information without full WM inference, (2) **Future-Aware Attention Masking** to preserve the original action-generation process and to block the queries from directly accessing the original non-visual tokens, and (3) **Action-Future Token Selection** to jointly exploit action-aware and future-aware importance to prune redundant and unimportant tokens. Extensive experiments across WAM backbones (i.e., RynnVLA-002 and UniVLA) in both simulation and real-world manipulation tasks demonstrate that FAVP achieves up to a 1.63× inference speedup on simulation benchmarks and a 1.67× speedup in real-world manipulation, with relatively small reductions in success rate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.