AC-Sparse: Action-Centric Dynamic Activation Sparsity for Efficient Vision-Language-Action Model Inference
Abstract
Vision-Language-Action (VLA) models enable robotic manipulation but require substantial computation during inference. Existing methods for efficient VLA inference often rely on coarse-grained layer skipping or token optimization, while computation reuse depends on temporal similarity between observations. Activation sparsification offers a fine-grained alternative by reducing linear computation while preserving the network structure and token sequence. However, existing sparsification criteria often overlook how perturbations at different tokens affect action generation, potentially degrading task performance at high sparsity levels. We propose AC-Sparse, a training-free dynamic activation sparsification method for VLA inference. Specifically, AC-Sparse combines attention weights from action representations with activation magnitudes and weight-column norms to estimate the cost of removing each token–channel activation. We formulate an attention-weighted expected squared local output perturbation objective and derive a separable upper bound that AC-Sparse minimizes under a fixed sparsity budget for each linear projection. We further prove that the derived sparsification criterion exactly minimizes the original objective when the weight columns are mutually orthogonal. Experiments across different VLA architectures and manipulation benchmarks demonstrate that AC-Sparse better preserves task success than the evaluated sparsification baselines at high sparsity levels, achieving an average success rate only 0.7 percentage points below dense inference on LIBERO at 65% activation sparsity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.