acceptodds
Under review as a conference paper at ICLR 2027

SGPrune: Subspace-Guided Token Pruning for Multi-View Vision-Language Models in Autonomous Driving

Abstract

Multi-view vision-language models (VLMs) support scene understanding and reasoning in autonomous driving, but processing visual tokens from multiple cameras incurs substantial computational and memory costs. Effective pruning must preserve task-relevant content and visual diversity while allocating a limited token budget across views. We propose SGPrune, a training-free framework for subspace-guided visual token pruning in multi-view VLMs. SGPrune combines three stages: Attention-Based Token Selection, PCA Subspace Construction, and Adaptive Multi-View Token Pruning. The first two stages run offline, using text-to-vision attention to select question-relevant visual tokens from a subset of the training data and summarize them into view-specific PCA subspaces. The third stage operates online during inference using the stored subspaces to estimate token relevance and guide diversity-aware selection before the multimodal projector. This selection is performed jointly across views under a shared budget, yielding input-adaptive view allocations. Experiments on autonomous driving benchmarks show that SGPrune achieves state-of-the-art performance among recent pruning methods, improving the DriveLMM-o1 Overall Reasoning score by 1.13 points and the MAPLM average score by 2.53 points over state-of-the-art baselines at about 90% token reduction. Under the same setting, SGPrune achieves a prefill speed on DriveLMM-o1 and a speed on DriveMM compared to the vanilla models, offering a superior balance between performance and efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.