Why does Post-Training Quantization Work?
Abstract
Post-training quantization (PTQ) compresses large language models (LLMs) by reducing weight precision and can largely preserve downstream task performance, even though they are never exposed to quantization error during pretraining. A common intuition is that low-bit formats such as NVFP4 induce only small errors in each weight matrix, so errors accumulate slowly through the model. However, we show that this explanation alone does not fully explain PTQ's empirical effectiveness. Specifically, when quantizing randomly initialized and pretrained models to the same precision, we observe that their weight quantization errors are nearly identical, yet randomly initialized models exhibit substantially larger hidden-state error than pretrained models. We attribute this discrepancy to a counteracting mechanism present in pretrained networks: the error newly introduced by a layer tends to oppose the error propagated from the preceding layer. These two components partially cancel, substantially slowing error accumulation across layers. Moreover, at the output layer, the probabilities of high-ranked tokens are more robust to the propagated errors. We quantify these mechanisms and validate our findings across multiple model families and quantization settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.