PiSO: Scale Optimization in Post-Training Quantization
Abstract
Post-training quantization (PTQ) compresses large language models by mapping weights to low-bit representations. While much work has focused on how weights are rounded onto the quantization grid, the scaling factor that defines the grid is typically chosen with simple, data-free heuristics. We observe that, under round-to-nearest (RTN) quantization, the grid assignments are piecewise constant in the scale, making the data-aware layer-output error piecewise quadratic over finitely many intervals, each with a closed-form minimizer. Building on this insight, we present PiSO (Piecewise Scale Optimization), an algorithm that leverages calibration data to efficiently compute the globally optimal channel-wise scales under RTN. We extend PiSO to group-wise quantization via tractable objectives. We also propose strategies for integrating it with error-correction methods such as GPTQ and Qronos, which greedily adjust weight assignments to compensate for the rounding errors of previously quantized weights. Experiments on Llama and Qwen models across multiple sizes and bit-widths show consistent improvements in perplexity and zero-shot accuracy, both standalone and combined with error correction, with gains increasing as the bit-width narrows. For instance, on Llama-3, PiSO reduces the 2-bit perplexity of Qronos by up to 7x, and at 3 bits, RTN with PiSO scales alone outperforms GPTQ with absmax scales.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.