Not Every Action Needs Every Token: Dynamic Visual Token Allocation for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models repeatedly execute the vision-language backbone at every control step, where visual tokens constitute a dominant portion of the inference cost. Existing visual-token pruning methods accelerate inference by applying a fixed retention ratio throughout deployment. However, we observe that sensitivity to sustained token reduction varies across tasks and control steps within a rollout. A uniform ratio can therefore miss opportunities to reduce computation while preserving task performance. To address this issue, we propose **D**ecision-importance-driven **V**isual-token **A**llocation (**DVA**), a dynamic token allocation framework that determines the retention ratio online based on the current observation and instruction. The key idea is to introduce **D**ecision **I**mportance (**DI**), a per-step supervision signal derived from counterfactual suffix rollbacks, which quantifies the downstream degradation caused by sustained token reduction starting at a given control step. A lightweight importance predictor learns to estimate DI from features already available in VLA models, and a monotone calibrator converts the predicted importance into the deployed retention ratio. DVA serves as a plug-in module for existing token pruners, introducing only about 1% per-step latency overhead while keeping both the pruning strategy and the policy parameters unchanged. Experiments show improved performance-efficiency trade-offs over fixed-ratio pruning. On OpenVLA-OFT, retaining 37% of visual tokens yields a success rate comparable to the full-token baseline at a 1.20× speedup, while retaining 22% maintains a 96.2% success rate at a 1.28× speedup. At matched average retention targets, DVA outperforms fixed-ratio pruning across three action heads, three pruning methods, and two benchmarks, demonstrating its effectiveness and generalization ability. Code and models will be released publicly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.