DART-VLA: Dynamic Action-Relevance Token Pruning for Efficient Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models map multimodal observations and instructions to continuous control actions. However, the large number of visual tokens imposes substantial computational overhead on the language-model backbone, hindering real-time closed-loop inference. Existing VLA pruning methods primarily rely on semantic attention or static importance estimates, leaving the potential of action signals for guiding token-level pruning largely underexplored. Our key insight is that a visual token can be better valued by its influence on the predicted action, as this directly reflects its contribution to the decision process. Based on this insight, we propose DART-VLA, a dynamic action-relevance token pruning framework for efficient vision-language-action models. DART-VLA backpropagates gradients from the action outputs to estimate each visual token's influence on action prediction and prunes tokens accordingly. DART-VLA first adapts the overall retention budget to the current motion state. Considering the distinct functional roles of the main and wrist views, we further allocate the token budget between them based on end-effector motion and independently retain the tokens most relevant to action prediction within each view. DART-VLA achieves a 97.40% average task success rate while delivering a 1.26\(\times\) inference speedup. Moreover, at a comparable inference speedup, DART-VLA outperforms the strongest pruning baseline by 1.90 % in average task success.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.