An -Law for Scale Multipliers in Low-Bit Post-Training Quantization
Abstract
Preserving language-model accuracy below four bits per weight motivates understanding how quantizers allocate their limited representable values. We ask whether post-training quantizers with different optimization procedures choose scales that follow a shared, predictive relationship. We compare their choices through a per-group multiplier that measures scale contraction relative to the group's largest weight magnitude. For AutoRound, GPTQ, and AWQ, its median follows , where is the largest positive integer quantization code and is fitted per method and configuration. A symmetric clipped mean-squared-error approximation interprets this contraction through the trade-off between clipping and rounding error. Across seven models from six families at two, three, and four bits per weight, the law achieves in-sample . Constants fitted on models with at most nine billion parameters predict the sampled median of Llama-3.1-70B with absolute errors below 0.016 for AutoRound and 0.008 for GPTQ at two and three bits per weight. Among the four tested methods, larger fitted corresponds to higher two-bit accuracy averaged across the seven models. For Llama-3.1-70B quantized to two bits per weight, AutoRound and its ablation without rounding iterations have similar sampled median scale multipliers yet retain 91% and 45% of the half-precision reference model's ten-task mean accuracy. The -law reveals shared structure in method-specific scale choices and supports prediction across bit widths and, for AutoRound and GPTQ, to a held-out 70B model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.