Beyond Action Deviation: Execution-Level Quantization Sensitivity for Vision-Language-Action Models
Abstract
Post-training quantization (PTQ) offers a practical way to reduce the deployment cost of vision-language-action (VLA) models. When uniform low precision cannot preserve control fidelity, mixed-precision PTQ relies on sensitivity estimation to identify which layers should retain higher precision. Existing sensitivity criteria mainly measure internal numerical distortion or action-level deviation, overlooking how quantization perturbations propagate through robot execution and influence task performance. Consequently, mixed-precision quantization guided by these criteria can lead to substantial performance degradation, particularly when applied to the sensitive action generation modules. We identify the need for an intermediate perspective between action deviations and task outcomes that captures robot-dependent execution information. To this end, we introduce Execution-Level Quantization Sensitivity (EQS), a layer-wise sensitivity criterion for mixed-precision quantization of VLA action generators. It models how quantization-altered action commands translate into horizon-level motion discrepancy and evaluates the discrepancy's execution significance under the current robot configuration. Together, these components provide an execution-aware sensitivity signal for layer prioritization. Experiments across simulation benchmarks and real-robot manipulation tasks show that EQS exceeds the best-performing baseline by 0.44 8.06 percentage points in average task success, with gains reaching 33.82 percentage points over individual baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.