Progress-aware Visual Token Pruning for Efficient Vision-Language-Action Manipulation
Abstract
Vision-language-action (VLA) models have shown great promise for embodied AI, yet repeatedly processing hundreds of visual tokens during robot execution incurs substantial computational overhead. Existing pruning methods reduce this cost by retaining task-relevant visual tokens and discarding irrelevant ones. However, the retention priority of task-relevant evidence changes with execution progress, particularly in long-horizon manipulation involving multiple subtasks. To address this issue, we introduce ProgPrune, a training-free method that adapts visual-token pruning to execution progress. Specifically, ProgPrune first parses instructions into ordered subtasks and uses execution feedback to track the active subtask and its execution phase. Guided by this progress, ProgPrune dynamically prioritizes visual tokens relevant to the active subtask over those associated with completed or future subtasks. This allows task-relevant but currently unnecessary tokens to be pruned, while full-instruction guidance preserves global context. As decoder representations evolve, we introduce Fresh-QK to recompute progress-aware token relevance at successive pruning layers. This keeps later pruning decisions from relying on outdated relevance estimates, without additional forward passes. Experiments on LIBERO, SimplerEnv, and CALVIN across multiple VLA architectures show consistent 1.20–1.44× speedups improvement with competitive task success rate. Notably, ProgPrune achieves particularly strong efficiency on long-horizon tasks (e.g., LIBERO-Long and CALVIN), demonstrating its effectiveness of progress modeling in visual token pruning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.