LPA-CACHE: Latency-Plateau Aware KV Cache Reuse for Efficient Vision-Language-Action Model Inference on GPUs
Abstract
Vision-language-action (VLA) models require low-latency inference for responsive robotic control. Previous works have shown that cross-frame reuse of visual key-value (KV) representations can reduce redundant computation. However, reductions in FLOPs do not necessarily translate into proportional latency reductions. Profiling the prefill-stage of VLA language-backbone on GPUs reveals a pronounced latency-plateau effect in dominant linear operators: reducing the number of recomputed tokens within a plateau yields marginal speedup, whereas crossing a plateau boundary yields substantial latency reduction. Motivated by this observation, we propose LPA-CACHE, a training-free method for accelerating VLA inference on GPUs through latency-plateau-aware KV cache reuse. A one-time profiling procedure identifies latency-plateau boundaries on the target GPU platform. These boundaries are combined with preselected reuse layers to construct a fixed-shaped reuse schedule. At inference time, visual tokens estimated to be task-relevant from previous-step attentions are protected from reuse, while the remaining visual tokens are ranked by inter-frame relative changes and progressively reused according to layerwise reuse budgets. Guided by observed cross-layer consistency in attention-based task relevance, LPA-CACHE extracts attention weights from only a small subset of layers, allowing FlashAttention-2 to be used in the remaining layers for further speedup. Extensive experiments across several representative VLA models on two robotic simulation platforms (LIBERO and CALVIN) demonstrate that LPA-CACHE achieves up to 1.38 language-backbone prefill speedup and up to 1.15 observation-to-action inference speedup with negligible loss in task success rate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.