HSA: Hierarchical Sensitivity-Aware Quantization for Resource-Constrained Embodied Models
Abstract
Embodied foundation models combine visual perception, language-conditioned reasoning, and action generation, but their memory footprint constrains deployment on resource-limited robotic platforms. Post-training weight quantization reduces this footprint, yet its perturbations can affect continuous action predictions and closed-loop execution. We propose Hierarchical Sensitivity-Aware Quantization (HSA), which augments a 4-bit AWQ backbone with BF16 residual blocks. An action-guided score combines a reconstruction-error proxy with action-output Jacobians, and a layer–channel–block hierarchy determines where the residuals are retained. Across LIBERO evaluations with VLA-Adapter, OpenVLA, and FastWAM, HSA improves average success over other methods by 2.55%–2.90% and remains within 0.30% of the original models. Meanwhile, it reduces GPU memory by 53.2%–67.2%. The quantized policy also matches the original policy in the real-robot setting under various conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.