SACQ: Structure-Aware Codebook Quantization for Ultra-Low-Bit Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models produce remarkably capable robot behavior, but their memory footprint makes deployment on resource-constrained robotic platforms challenging. Weight quantization is a natural remedy, yet post-training quantization (PTQ) recipes inherited from large language models degrade sharply when applied to VLAs at ultra-low precision. Two limitations become particularly important: existing low-bit PTQ methods typically assess quantization error at the layer or projection level without modeling the downstream operation through which it propagates, and allocate precision without directly measuring how reducing a layer’s precision changes the policy action. We introduce SACQ, an ultra-low-bit weight-only vector-codebook quantization framework for VLAs. SACQ combines (1) structure-aware codebook fitting, which fits each layer using a role-specific quadratic metric derived from the operation its projection feeds; (2) action-aware mixed-bit allocation, which assigns precision according to decoded-action degradation as individual layers are compressed; and (3) two-stage post-quantization refinement, which first reassigns codeword indices with second-order error compensation and then calibrates codebooks and row-group scales against the decoded action. On GR00T N1.7, SACQ retains 93.85% average success across all four LIBERO suites at approximately 1.5 bits per language/action weight, only 0.25 percentage points below the 94.10% BF16 reference, while reducing resident model memory from 5.677 GB to 1.175 GB, a 4.83× reduction. On SimplER Google Robot, at approximately 2 bits per language/action weight, SACQ achieves 63.76% and 52.27% success under Visual Matching and Variant Aggregation, respectively, compared with 58.90% and 53.45% for BF16.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.