acceptodds
Under review as a conference paper at ICLR 2027

Seeing Action-Critical Evidence: Token Pruning with Internal Action Priors for Vision-Language-Action Models

Abstract

Vision-Language-Action (VLA) models translate visual observations and language instructions into continuous robotic actions. However, their dense visual token sequences introduce substantial decoding overhead, hindering efficient real-world deployment. Recent trainable token pruning methods reduce this cost, yet their pruning criteria remain largely action-agnostic: they often retain task-irrelevant background tokens while discarding fine-grained visual cues essential for precise control. In this work, we investigate whether pretrained VLA models internally encode action-aware signals for token selection. Our analysis suggests that: (i) action-critical visual tokens show an association with value vectors in deep Feed-Forward Network (FFN) layers; and (ii) these value vectors define actionable residual subspaces whose projection energy approximates local action saliency. Motivated by these insights, we propose ActPruner, a novel Action-Anchored Dynamic Token Pruning approach, which constructs an action-aware pruning metric from intrinsic action priors encoded in deep FFN layers. Technically, ActPruner constructs an Action Information Extraction Module (AIEM) with three progressive steps: Deep Prior Extraction to extract compact action anchors from deep FFN value vectors, Semantic Space Adaptation to align them to shallow pruning layers, and Instruction Modulation to generate action-conditioned pruning queries. This enables dynamic preservation of manipulation-relevant affordance regions while suppressing redundant background tokens. Extensive experiments on VLA-based robotic manipulation benchmarks verify the superiority of our ActPruner over state-of-the-art VLA pruning methods. Code is available in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.