When Does the Calibration Source Matter for Low-Bit GPTQ Quantization of LLMs? A Seed-Paired Study
Abstract
Earlier studies found that the calibration source has little effect on four-bit weight quantization of current language models, while a three-bit study of one reasoning model found GPTQ highly sensitive to the calibration domain. We measure both precisions on the same calibration draws. Qwen3-8B, Llama-3.1-8B and Mistral-7B-v0.3 are quantized with weight-only GPTQ at three and four bits using WikiText, random-token or prompted model-generated calibration text, which we call Diverse, with five calibration draws shared across sources and precisions and a budget of tokens, and are evaluated on WikiText-2 perplexity and all 14,042 MMLU questions. On the recorded MMLU readout, the Diverse-minus-WikiText contrast is , and points at four bits and , and points at three bits, every three-bit seed-paired 95% interval lying above zero, while WikiText has the lower perplexity in all 15 model–seed pairs; the source-by-precision interaction averages points, with a 95% interval of over seed blocks. Choosing the source by WikiText perplexity thus forgoes , and MMLU points at three bits relative to the best source in hindsight and under one point at four bits, and choosing it by accuracy on held-out questions lowers this three-bit regret for Qwen3-8B and Llama-3.1-8B. Three-bit calibration draws vary – times more than four-bit draws. Registered replacement runs that score every prompt untruncated reproduce the three-bit gain in every seed of all three models, and on 15 fresh Qwen3-8B draws in three registered batches the gain and the interaction average and points with intervals above zero, as they do on ARC-Challenge for the ten of these draws scored on it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.