acceptodds
Under review as a conference paper at ICLR 2027

Beyond Scalar Recovery: Residual Geometry in Quantized Vision-Language-Action Models

Abstract

Post-quantization recovery in vision-language-action (VLA) models is typically summarized by scalar error or task-level metrics, which measure how much error remains but not how that error is organized across a multi-step action chunk. We characterize this missing structure through teacher-relative residual geometry in a controlled OpenVLA-OFT setting. Across 13 heterogeneous ternary quantization configurations, mean pairwise Spearman similarity increases from 0.037 before adaptation to 0.557 afterward, with a demonstration-bootstrap median gain of +0.520 (95% CI [0.472, 0.568]). This geometric reorganization is not explained by uniform error shrinkage, and scalar recovery magnitude can diverge substantially from final residual alignment. The shared-alignment trend persists across additional optimization seeds, matched ground-truth-only adaptation, and a distinct 3-bit quantizer. Matched full-precision controls further demonstrate that shared alignment does not require an initial quantization perturbation. Finally, a leave-one-out template formed directly from other configurations’ adapted residuals strongly predicts the final residual geometry of a held-out configuration, reaching mean Spearman correlation 0.678 and substantially outperforming both raw geometry and a leave-one-out recovery-field predictor. Shared chunk-position and action-dimension effects explain much of this predictability, while the complete template retains additional predictive structure in reconstruction error.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.