OPTQ: Quantization Error is Relocated, Not Eliminated
Abstract
Post-Training Quantization (PTQ) inherently cannot eliminate the quantization error it generates; it can only determine where to distribute this error. The magnitude of the residual injected during rounding to a fixed grid is strictly determined by the grid itself, remaining identical across any algorithm. Consequently, what a quantizer actually controls is not the magnitude of the error, but the allocation of the error across channels, basis directions, and time. This paper formalizes this perspective. For an arbitrary linear operator receiving a quantized tensor, the distortion surviving to the output is dominated by the operator's spectrum. Thus, a good quantizer is one that steers the error toward the directions that the operator attenuates. This framework not only designates operators for each axis but also, equally importantly, predicts which axes are not worth optimizing. For instance, the time axis of the Value cache is dismissed prior to any empirical experiments, solely based on the observation that the measured lag-1 attention autocorrelation is approximately 0.06. By instantiating this framework, we derive a calibration-free sequential quantizer for activations, where the downstream operator is not estimated from data but exactly determined by the layer's weights. For the Key cache, we formulate an operator combining channel mean subtraction and per-head Hadamard rotation. Empirically, under a W4A4KV4 setting, our training-free configuration achieves a perplexity of 5.85 on Llama-2-7B, outperforming recent methods that rely on learned transformations (QuaRot, FlatQuant, SpinQuant, and OSTQuant). Incorporating transformations learned to pass through the relocation operator establishes a new state-of-the-art, yielding a perplexity of 5.714 on Llama-2-7B and 6.899 on Llama-3-8B.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.