From Jump Degradation to Warped-Space Quantization
Abstract
The growing parameter scale of large language models imposes substantial memory and computational costs, making weight quantization an important approach to efficient deployment. Quantization represents continuous weights with a limited set of discrete levels, and its effectiveness depends increasingly on how this limited precision is allocated as the bit width decreases. Existing methods improve this allocation through optimized rounding, scaling, codebooks, or transforms, but they are commonly designed using static statistics collected from weights or calibration data. During model execution, however, quantization error evolves dynamically: fixed weight residuals interact with changing activations that may already contain perturbations propagated from upstream layers. To characterize this process, we progressively replace full-precision weights with their quantized counterparts and compare the resulting intermediate activations with those of the original model. We observe jump degradation, in which activation error rises abruptly at a small number of layers and continues to amplify downstream. Further analysis identifies the mismatch between equally spaced quantization levels and non-uniform weight distributions as one controllable source of the residual involved in this process. We therefore propose Warped-Space Quantization (WSQ), which maps weights into a flatter space for uniform sampling and applies the inverse map to construct distribution-matched non-uniform levels. By allocating more of the limited precision to high-density regions, WSQ reduces the inputindependent weight residual before it interacts with changing activations. Experiments across multiple models and benchmarks show that WSQ consistently improves quantized-model quality and outperforms the compared baselines under directly comparable low-bit quantization settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.