acceptodds
Under review as a conference paper at ICLR 2027

QFast: Fast Quantized Inference for Robotics Vision-Language-Action Models

Abstract

Vision-Language-Action (VLA) models have demonstrated strong performance in robotic manipulation by unifying visual perception, language understanding, and action prediction within transformer-based architectures. However, their deployment on resource-constrained edge devices faces critical challenges: inference latency and memory requirements that prevent real-time operation. Existing quantization methods prioritize accuracy preservation and memory compression while largely neglecting inference speed—a bottleneck for real-time robotic control requiring sub-100ms latency. We present QFast, a quantization inference system designed to accelerate VLA models on edge devices. QFast employs speed-oriented quantization that balances accuracy and latency through module-specific bit-width selection, per-token weight quantization with dynamic per-channel activation quantization, and selective preservation of precision-sensitive layers. To exploit hardware capabilities, we implement optimized GEMM kernels tailored to diverse VLA matrix dimensions, enabling native low-bit computation on Tensor Cores. Evaluated on Pi0.5 and OpenVLA, QFast achieves 1.3 end-to-end speedup with up to 3.6 memory reduction while maintaining task performance on LIBERO benchmarks. QFast represents the first quantization system enabling hardware-accelerated low-bit inference for VLA models, making real-time edge deployment practical. Code is available at https://anonymous.4open.science/r/QFast-D157/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.