SparseVLA: Towards Smaller And Faster Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models offer a promising approach to general-purpose robotic control, but their large parameter counts contribute to substantial storage requirements and inference latency, hindering efficient deployment. Weight pruning is a representative approach to addressing this problem. However, when applied to VLA models, existing weight pruning methods typically use information from each control step in isolation, overlooking the sequential nature of VLA control. In practice, two challenges arise: 1) Methods may reconstruct individual outputs well but fail to preserve how they change across steps; 2) Sparse execution can slow down inference when runtime overheads outweigh computational savings. To solve these challenges, we introduce SparseVLA, a training-free framework for hardware-efficient VLA weight pruning and deployment. We find that activation changes between steps reflect task progression, suggesting their potential as a complementary signal for weight pruning. Since preserving inter-step activation changes can favor different pruning patterns than reconstructing individual outputs, SparseVLA balances the two criteria when selecting and reconstructing weights. To turn this sparsity into practical inference gains, SparseVLA combines shape-aware sparse execution with operator fusion, reducing memory traffic and kernel launch overhead. Across LIBERO, SimplerEnv, and real-world manipulation tasks, SparseVLA maintains or improves aggregate task success relative to dense baselines on two VLA backbones, and GR00T, while reducing model footprint and inference latency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.