Scale Sensitivity in Low-Bit Post-Training Quantization: Curvature of the Quantization Error Landscape
Abstract
Post-training quantization requires choosing a grid scale, and max-based choices can fail sharply at low bit-widths. For i.i.d. Gaussian weights and activations with sufficiently large effective rank, we prove uniform convergence of the normalized round-to-nearest loss to a Gaussian quantization objective. Its unique optimal scale decreases with the number of levels, while its relative-scale curvature decays approximately exponentially with bit-width. We verify the rank condition for wide random MLPs. Across five LLMs, GPTQ scale sensitivity likewise falls with bit-width: scale choice matters greatly at 2–3 bits but little from 6 bits onward. The Gaussian-optimal scale performs poorly on raw weights; after Hadamard incoherence processing, it matches the best searched rule from 3 bits onward but remains worse at 2 bits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.