acceptodds
Under review as a conference paper at ICLR 2027

Policy-Functional Token Pruning for Efficient Vision-Language-Action Models

Abstract

Visual token pruning can accelerate vision-language-action (VLA) inference, but accurate importance rankings do not necessarily yield an effective retained set. Our token-deletion diagnostics show that attention closely tracks how much removing a token changes an intermediate action-related representation, yet different influential tokens can change it in similar directions. In light of these observations, we introduce Policy-Functional Token Pruning (PFT-PRUNE), a training-free method that considers both individual influence and pairwise directional alignment. Specifically, at an early attention layer, we characterize each visual token by the local sensitivity of an action-related representation to its key, and use shared random projections to construct a compact policy-functional fingerprint (PFF). PFF energy provides the scalar priority, while fingerprint alignment identifies potential functional redundancy; Functional NMS uses both to select tokens under an exact budget before physical pruning. Across four LIBERO suites, PFT-PRUNE achieves 94.1% and 93.2% average success on and OpenVLA-OFT at 20% nominal visual-token retention, with inference speedups of and , respectively. On four physical-robot tasks, it achieves a speedup at 50% nominal retention while preserving 95.9% of 's average success rate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.