MemCache: Memory-Augmented Adaptive Token Caching for Efficient Vision-Language-Action Inference
Abstract
Vision-Language-Action (VLA) policies facilitate generalized robotic manipulation by mapping visual scenes and task instructions to low-level robotic actions. Nevertheless, their heavy computations fail to meet the stringent real-time control constraints and limited computational budgets of physical robots. As an efficient inference paradigm, VLA caching alleviates computational burdens by exploiting temporal redundancy across consecutive robotic observations and reusing static token states. However, existing methods exhibit several limitations: their caching decisions generally depend on hand‑crafted proxy signals often misaligned with control policy demands, and they rely only on adjacent frames without exploiting longer-term temporal context. To this end, we present MemCache, a novel learnable memory-augmented caching framework for high-efficiency and reliable VLA inference. Specifically, we leverage learnable query banks to generate token‑level importance scores. This eliminates dependence on hand‑crafted proxy supervision and externally defined signals, enabling the token‑selection strategy to be fully optimized with the control objective. Additionally, we integrate a recurrent visual memory state to capture long-range temporal dependencies in manipulation behaviors, providing history-aware corrections for caching decisions. We further develop a linear-warmup training schedule to coordinate the above two components and facilitate joint optimization. Equipped with these merits, MemCache delivers reliable and efficient VLA inference for robotic manipulation tasks, e.g., higher task performance with speedups over VLA caching baselines on the LIBERO benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.