MARK-VLA: Mixed-Precision Action-Risk-Guided Quantization with Mixed-Weight Kernels for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models couple visual perception, language reasoning, and action generation within a unified policy, but their large model size and computationally intensive inference impose substantial memory and latency overheads. We present MARK-VLA, a post-training quantization framework that enables efficient VLA inference through channel-wise mixed-precision quantization and tailored GPU kernel support for mixed-precision weights. Instead of pursuing extreme compression for individual modules, MARK-VLA applies W2/W4/W8 weight quantization with A8 activation quantization across the major perception-to-action stages to reduce the overall memory footprint. We propose a per-module precision allocation method that automatically determines channel-wise precision based on gradient-based action risk evaluated under a quantized reference policy. To account for the discrepancy between the risk estimated for individual quantization decisions and their combined effect, we adopt a progressive precision reduction scheme that gradually lowers the quantization precision. Finally, to translate these fine-grained precision assignments into practical acceleration, we develop fused mixed-weight GPU kernels together with trace-guided per-layer specialization to optimize the complete execution path of quantized Linear layers. Experiments on LIBERO across six VLA policies and three GPU platforms demonstrate up to compression and up to end-to-end speedups, while the largest decrease in average success rate is only 0.82 percentage points. These results demonstrate that action-aware precision allocation and deployment-oriented kernel co-design can jointly achieve compact and efficient VLA inference with minimal accuracy degradation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.