acceptodds
Under review as a conference paper at ICLR 2027

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement

Abstract

Vision-Language-Action (VLA) models is a promising approach for generalizable robotic manipulation. Recently, VLA models that employ flow-matching-based action heads have shown remarkable success, as they can generate precise and smooth action sequences while capturing multimodal action distributions. However, these action heads require iterative denoising, which acts as a major computational bottleneck in inference, posing a critical challenge for real-time VLA deployment. To address this challenge, we propose ActionCache, a plug-and-play external cache that opportunistically reuses past intermediate actions. Specifically, ActionCache stores intermediate actions with compact multimodal keys and retrieves them from similar past contexts to warm-start generation from the vicinity of target actions. This mechanism reduces the number of required denoising steps, leading to substantially lower inference latency. Experimental results in simulation and real-world environments demonstrate that ActionCache maintains high task success rates in a low-latency regime, achieving action head inference acceleration of up to and for representative flow-matching-based VLA, and GR00T-N1.6, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.