Grounding Attention in Motion: Proprioception-Guided Visual Token Pruning with Adaptive Layer Selection for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models benefit from the perceptual generalization of pretrained vision-language backbones, but repeatedly processing hundreds of visual tokens makes closed-loop inference costly. Visual token pruning can reduce this cost, yet existing methods typically rank tokens using only vision-language interactions and prune at a fixed backbone layer, overlooking both the robot's current state and the input-dependent stability of token rankings. We introduce PropPrune, a training-free framework that adapts both which visual tokens are retained and where they are pruned. A proprio-guided importance score uses each instruction token's attention to the proprioceptive state to weight its attention to visual tokens, grounding visual relevance in the current embodied context. An adaptive layer selector then measures the Jaccard similarity between top-k visual-token sets from adjacent shallow layers and defers pruning to a later layer when the shallow ranking is unstable. Both components use attention maps already produced by the pretrained VLA, requiring no fine-tuning, auxiliary model, or additional forward pass. On LIBERO, PropPrune achieves the highest average success rate among the evaluated pruning baselines at both 50% and 25% visual-token retention, remaining within 0.7 and 1.3 points of the full-token model while reducing latency by 16.0% and 34.9% and FLOPs by 38.1% and 53.6%, respectively. On LIBERO-Plus, it improves over FastV under seven controlled distribution shifts and remains within 0.26 points of the full-token reference. On two real-world tasks, 50% retention matches full-token success while reducing latency by 19.1%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.