Exploring What to Preserve for Reasoning and Action in Structured Pruning of Driving VLAs
Abstract
Reasoning vision–language–action (VLA) models for autonomous driving use chain-of-thought (CoT) reasoning to guide action generation, but incur substantial memory and computational costs. We study structured pruning for these models, focusing on preserving both scene understanding and trajectory prediction. Our analysis of Taylor importance in Alpamayo-1.5-10B shows that reasoning and trajectory objectives assign different importance to the same vision–language model (VLM) structures. In the action expert, MLP channels exhibit more stable importance rankings across denoising steps than query heads, and a small fraction of channels accounts for most of the measured importance. Based on these findings, we propose , which combines importance ranks from both objectives for VLM pruning and aggregates stepwise importance to select a fixed subset of expert MLP channels. At 24.0% total parameter reduction, VLM-only pruning with retains 94.0% of the dense model's LingoQA score and achieves the highest closed-loop scene scores among methods at the same budget. Removing 93.75% of expert MLP channels increases total parameter reduction to 39.4%, with nearly unchanged open-loop errors relative to VLM-only pruning and competitive closed-loop performance. The resulting model achieves a end-to-end inference speedup and 36.6% lower peak memory than the dense model, without post-pruning fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.