DepthPruner: Visual Token Pruning with Pluggable Depth Guidance for Efficient Vision-Language-Action Inference
Abstract
Vision-language-action (VLA) models use visual observations and language instructions to generate robot actions, but processing visual tokens is computationally expensive. Attention-based visual token pruning reduces this cost by retaining a subset of visual tokens. However, we observe that some potentially task-irrelevant tokens can receive high attention scores and compete with task-relevant tokens for a limited token budget. Depth provides geometric cues that complement attention scores for token selection. Motivated by this observation, we propose DepthPruner, a pluggable, training-free method that augments attention-based visual token pruning methods with depth guidance. It expands the candidate set provided by a base pruner and uses mean depth, residual variance, and neighborhood contrast to select tokens within the original budget. Experiments on OpenVLA-OFT across four LIBERO suites show that DepthPruner improves the average task success rate of all three base pruners at every tested token retention ratio. By replacing at most 8 of FastV's selected tokens per view, DepthPruner increases the success rate on LIBERO-Object by 9.0 percentage points at 25% token retention. Timing measurements show that depth refinement adds little inference overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.