CARE-VLA: Capacity Allocation via Response Evaluation for 1-Bit VLA Quantization
Abstract
Vision-language-action (VLA) models are rapidly advancing embodied intelligence, but their large parameter footprints remain a major obstacle to deployment on resource-constrained robotic platforms. One-bit post-training quantization (PTQ) offers an attractive route to extreme model compression, yet preserving model behavior at this precision remains challenging. To the best of our knowledge, vector quantization has not yet been explored for VLA models in the 1-bit regime, where deciding which concrete reconstruction deserves scarce additional capacity becomes critical. We introduce Capacity Allocation via Response Evaluation (CARE-VLA), a post-training vector-quantization framework that allocates capacity among concrete low-bit reconstruction candidates according to their response-level distortion. Rademacher-projected reverse-mode pullbacks efficiently estimate native-response sensitivity without materializing full Jacobians, allowing concrete reconstruction errors to be evaluated directly in response space. Complete-policy component interventions further calibrate heterogeneous native-response gains onto a common action-discrepancy scale, enabling reconstruction candidates across vision, language, action, and auxiliary modules to compete under a shared bit budget. At 1.09 bits per weight (BPW), CARE-VLA achieves 91.0% and 88.9% mean LIBERO success on and OpenVLA-OFT, respectively, and 69.7% visual-matching and 53.2% variant-aggregation success on SIMPLER with CogACT. Executed linear weights achieve storage compression relative to 16-bit storage, including codec overhead. On the RTX 4090, the quantized model achieves a measured speedup. Our work advances extreme VLA compression toward practical deployment on resource-constrained robotic hardware.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.