Redundancy or Relevance: What Language-Guided Token Pruning Selects in Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) model token pruning methods rank tokens typically by computing an importance score based on attention weights, spatial redundancy via inter-token similarity, or task relevance via language conditioning. In this work, we analyze pre-decoder language-vision representations of VLA models. We start this paper by proposing a simple, yet effective, language-vision cosine-similarity based pruning method. Surprisingly, the simple pruning method outperformed more complex state-of-the-art pruning methods. So to better understand how pruning works, we study in more details how cross-modal similarities drive pruning. We find that in many cases pruning is effective in VLAs, not because of semantic grounding, but rather due to the criterion's projection geometry acting as a redundancy suppressor along the visual typicality axis. Since each backbone produces a uniquely structured embedding space, we characterize its geometry to determine what the score actually ranks. We show that when the geometry allows, this signal could be sufficient to determine which tokens to drop, despite the lack of strong semantic alignment. We test over four different VLA models with three different robotic datasets, showing how our language pruning method and the simpler method achieve high accuracy with the lowest latency. For example, on OpenVLA-OFT, we achieve up to latency reduction compared to vanilla OpenVLA-OFT, with only 1% drop in success rate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.