acceptodds
Under review as a conference paper at ICLR 2027

Beyond Shadow Weights: Quantization-Aware Training as Quantized-Endpoint Descent

Abstract

Quantization-aware training (QAT) updates a full-precision shadow weight but deploys the quantized endpoint . Existing explanations for QAT largely view its success through the lens of shadow weights: QAT can move toward flatter basins, gain robustness from oscillations, or balance the shadow loss against the quantization error . These perspectives do not directly explain the empirical observation that the deployed endpoint loss improves even as the shadow loss increases substantially. In this paper, we offer a different explanation by treating QAT as finite-grid endpoint dynamics. Motivated by the approximate normality of rescaled pretrained weights, we propose an idealized residual-phase model that yields a crossing law that determines which coordinates cross quantization boundaries after a shadow update. We also find that signal imbalance creates asymmetric oscillations that favor weaker-gradient endpoints. This motivates QAR (Quantization with Amplified Routing), an algorithmic framework that directly operates on the quantization code. We further prove feasible-gradient guarantees for QAR with a family of power amplifier. Post-training experiments on three large language models support the endpoint view and show that QAR can achieve performance comparable to QAT with up to 24.1% peak memory saving.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.