acceptodds
Under review as a conference paper at ICLR 2027

QCache-VLA: Quantization and Diffusion Cache for Efficient Vision-Language-Action Policy Inference

Abstract

Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robot control. But their deployment on resource-constrained robotic platforms remains challenging due to the heavy Vision-Language Model (VLM) backbone and multi-step diffusion-based action generation. Existing acceleration methods are not directly suited to VLA diffusion policies: post-training quantization can suffer from image-text distribution mismatch and action-agnostic quantization decisions, while fixed-interval diffusion caching cannot adapt to non-uniform redundancy across denoising steps. To address these issues, we propose QCache-VLA, a discrepancy-aware acceleration framework for efficient VLA inference. For the VLM backbone, we introduce action-aware post-training quantization with modality-aligned smoothing candidate generation, action-aware candidate selection, and modality-balanced weight reconstruction. For the action expert, we propose a residual-calibrated dynamic diffusion cache that reuses deep residuals only when adjacent denoising steps are sufficiently similar. We further integrate INT8 Triton kernels and diffusion-step precomputation to enable practical end-to-end acceleration. Experiments in simulation and real-world tasks show that QCache-VLA preserves task performance while reducing GPU memory usage by and achieving a end-to-end speedup.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.