acceptodds
Under review as a conference paper at ICLR 2027

Divide and Prune: Geometry-Driven Token Pruning for Efficient VLA Driving Models

Abstract

Vision-Language-Action (VLA) models have emerged as a premier paradigm for embodied AI, enabling scene comprehension and auto-regressive action generation in autonomous driving, yet their deployment in the real-world system is hindered by the excessive number of visual tokens from high-resolution inputs. Previous general visual token pruning approaches primarily exploit text-visual attention, visual attention maps or token similarity, inherently blinding themselves to the non-uniform space and risk-asymmetric nature of driving environments. In this work, we propose Divide and Prune (DiP), a geometry-driven visual token pruning strategy for VLA driving models. DiP utilizes the depth estimation as spatial priors to divide the visual field into near-field and far-field regions. For near-field tokens, DiP fuses the edge energy with the token informativeness to retain local details critical for obstacle avoidance. For far-field tokens, we propose the depth-modulated diverse token selection to maximize semantic diversity while mitigating the semantic dilution caused by spatial overlap in 2D images. Experiments on the nuScenes and the Bench2Drive benchmarks show that DiP achieves strong performance while delivering an approximately 3.4 inference speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.