Not All Calibration Data Are Equal: Optimizing Calibration Data for Post-Training Quantization
Abstract
An often overlooked source of performance variation in post-training quantization (PTQ) is the calibration data used to estimate quantization-related statistics. Even under the same model and quantization configuration, different calibration data can lead to substantially different performance. We propose CalibSelect, a calibration data assessment and selection framework based on model responses. CalibSelect introduces two metrics: the calibration activation range score (CARS), which measures the spread of activation responses while mitigating the influence of extreme samples, and the layer sensitivity discrimination score (LSDS), which measures how clearly and stably calibration data reveal differences in quantization sensitivity across layers. Based on these metrics, CalibSelect adopts a two-stage selection strategy that first identifies several promising calibration data sets from the candidate data pool and then refines them through indicator-guided local optimization. CalibSelect is compatible with different PTQ methods and does not modify their underlying algorithms. Experiments on Qwen2.5-7B and LLaMA3.1-8B using AWQ and AutoRound show that CalibSelect consistently outperforms random sampling and existing selection methods. The largest reduction in quantized model perplexity reaches 8.23%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.