acceptodds
Under review as a conference paper at ICLR 2027

VADP-VLA: View-Adaptive Dynamic Visual Token Pruning for Efficient VLA Inference

Abstract

Vision-language-action (VLA) models generate robot actions from visual observations, language instructions, and robot states, but the large number of visual tokens produced by multiple cameras increases inference cost. Existing methods reduce visual computation by exploiting visual saliency, temporal redundancy, or action relevance. Because camera views differ in both content and temporal dynamics, however, the validity of historical information and the appropriate visual token budget must be assessed separately for each view. We propose VADP-VLA, a training-free framework for view-adaptive dynamic visual token pruning. It estimates language-conditioned visual token saliency within each view and transfers historical saliency through a per-view keyframe cache. Agreement between current and historical saliency controls cache refresh and token retention independently for each view. Tokens that are not retained are removed before later language-model layers, shortening the sequence processed by subsequent computation. Across the four LIBERO suites, VADP-VLA achieves an average success rate of 95.25%, compared with 94.55% for the unpruned OpenVLA-OFT baseline, while reducing normalized language-model FLOPs to 41% of the baseline and model-internal latency from 73.12 ms to 48.89 ms, a 1.50× speedup. In LIBERO experiments with , average success changes from 96.30% with the unpruned baseline to 94.65%, normalized language-model FLOPs fall to 45%, and model-internal latency decreases from 37.38 ms to 25.43 ms. Across four real-robot tasks, each method is evaluated in 80 trials: VADP-VLA succeeds in 63/80 trials versus 67/80 for the unpruned baseline, while VLA model-call latency falls from 59.93 ms to 42.58 ms, approximately a 1.41× speedup. These results show reduced VLA inference cost in the evaluated settings, alongside a modest decrease in real-robot task success relative to the unpruned baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.