acceptodds
Under review as a conference paper at ICLR 2027

SP-QAT: Spectral-Preserving Quantization-Aware Training for Extreme Low-Bit Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models are rapidly advancing and becoming central to embodied AI, yet their deployment on edge devices remains challenging. Since in VLA model, LLM backbone dominates memory consumption and inference latency, aggressive quantization of this LLM module presents a straightforward solution. Unfortunately, existing VLA quantization techniques mostly operate at 4-bit or higher bit widths and avoid fixed sub-4-bit quantization because of its catastrophic performance collapse, leaving it largely underexplored. To bridge this gap, we propose SP-QAT, an efficient quantization-aware training framework (QAT) that enables fixed 2-bit quantization of the Large Language Model (LLM) backbone in VLA, to our knowledge, for the first time. Through theoretical and experimental analysis, we identify that preserving the cognitive feature spectral subspace is the key to extreme low-bit quantization for the LLM backbone of VLAs. Therefore, we first constrain the weight-update subspace to preserve the pretrained model priors. Second, we propose Soft-Mask Gradients with Severity-Aware Scale Regulation to mitigate outlier-driven step size expansion from causing quantized weights to collapse into noise. Finally, we propose Difficulty-Aware Curriculum Learning with SNR-Guided Static Reweighting to emphasize hard samples and low-noise timesteps, thereby preserving the action-sensitive energy in the cognition feature spectrum. Extensive experiments on SIMPLER with CogACT and show SP-QAT only results in a small drop of compared to FP16 baselines, with 75%–84% memory savings and inference speedup. Code is available as supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.