DriveVLA-Cache: Training-Free Acceleration of Vision-Language-Action Models for Autonomous Driving
Abstract
Reasoning vision-language-action (VLA) models show strong performance in autonomous driving, but inference latency makes real-time deployment challenging. The cost spans four stages: visual encoding repeatedly processes overlapping camera frames; multimodal prefill recomputes shared context; autoregressive decoding often regenerates identical reasoning text; and trajectory generation repeatedly evaluates the action network over successive integration steps. Reasoning decoding contributes the largest share in our profiling, but accelerating it alone leaves substantial computation in the remaining stages. We propose DRIVEVLA-CACHE, a training-free framework that integrates Reasoning Recontextualization with visual feature, prompt K/V, and velocity-field caching to accelerate the full VLA inference pipeline. Visual feature caching avoids re-encoding overlapping frames, while selective prompt key-value (K/V) reuse reduces repeated context computation. For reasoning, the central challenge is that reusable text has cached K/V states tied to earlier observations. Our Reasoning Recontextualization periodically re-encodes retained reasoning text under the current observations in one forward pass, updating the states used for action generation without repeating autoregressive decoding. Finally, velocity-field caching reuses or extrapolates recent fields between selected action-network evaluations, reducing computation while preserving all integration updates. Together, these mechanisms accelerate the pipeline while producing a new trajectory every planning cycle. On Alpamayo 1.5-10B across 913 closed-loop ALPASIM scenes, DRIVEVLA-CACHE achieves the highest mean driving score and lowest latency among evaluated methods. It delivers 3.5× driver-inference and 2.4× end-to-end speedups over the full-computation baseline, while increasing mean AlpaSim Score from 70.8 to 73.4.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.