The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization
Abstract
Post-training quantization (PTQ) methods typically quantize a pre-trained large language model (LLM) sequentially using a small set of calibration samples, achieving high data and computational efficiency. Specifically, the model is partitioned into a series of units, such as transformer blocks, with one unit quantized at each stage. State-of-the-art PTQ methods are predominantly learning-based, optimizing auxiliary quantization parameters (e.g., scaling factors, rotation matrices, clipping thresholds, and adapters) via gradient descent to minimize a reconstruction loss, thereby mitigating the impact of outliers and reducing quantization errors. In learning-based PTQ, a common practice is to use mean squared error (MSE) as the reconstruction loss function, yet its induced optimization behavior remains largely unexplored. In this work, we take a holistic view of sequential quantization and systematically investigate how optimization evolves from the first quantization stage to the last, aiming for a deep understanding of optimization in learning-based PTQ schemes. Through extensive empirical studies spanning representative learning-based PTQ methods, model families, model scales, architectures, and quantization settings, we consistently uncover *Optimization Imbalance*: reconstruction loss magnitudes vary dramatically across stages, accompanied by highly uneven gradient magnitudes and parameter updates under MSE. We term the cross-stage range of loss magnitudes the *reconstruction loss scale*. Our analysis reveals that MSE translates the unexpectedly large reconstruction loss scale into highly uneven gradient magnitudes, which in turn lead to uneven optimization strength across quantization stages. This finding suggests a general principle for improving learning-based PTQ: optimization strength across stages should be decoupled from the reconstruction loss scale. Theoretically, we show that root mean squared error (RMSE) naturally realizes this principle through implicit gradient normalization, thereby enabling balanced optimization across quantization stages. Moreover, RMSE can be naturally defined over different feature groupings, yielding four variants at the sample, channel, token, and element levels. As a drop-in replacement for MSE, RMSE consistently improves quantization performance across language modeling, commonsense reasoning, mathematical reasoning, and coding tasks, with all four variants outperforming MSE and the gains becoming increasingly pronounced under more aggressive quantization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.