AIA-VLA: Adaptive Inference Acceleration for Vision-Language-Action Models
Abstract
Vision-language-action models (VLAs) combine the semantic understanding of pretrained vision-language models with diffusion-based action generation, demonstrating strong capabilities in robotic manipulation. However, their computational cost and inference latency hinder deployment on resource-constrained robotic platforms. Acceleration strategies with fixed computational budgets cannot readily accommodate varying computational demands across observations and inference stages, potentially compromising task performance. We present AIA-VLA, a training-free adaptive inference acceleration framework that uses intermediate inference signals to jointly regulate visual token budgets, language model execution depth, and feature reuse in diffusion-based action models. Specifically, AIA-VLA determines visual token budgets through cumulative task-relevant attention, uses recent inter-layer representation stability to trigger subsequent layer skipping, and uses feature drift between consecutive fully computed diffusion steps to decide whether to reuse attention and MLP branch outputs in the next step. These modules can operate independently or in combination, enabling input-dependent computation allocation without modifying model parameters. Evaluated on the SIMPLER benchmark with CogACT as the base model, AIA-VLA reduces floating-point operations by 50% and achieves a 1.54× inference speedup, while improving task success rate by 1.16 percentage points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.