acceptodds
Under review as a conference paper at ICLR 2027

Which Language Pays for Quantisation? The Calibration Corpus under GPTQ

Abstract

Open language models are evaluated at full precision, but people often run 4-bit copies quantised by third parties. We audit the 10 most-downloaded quantised releases of Qwen3-8B by the shift in the base model's confident answers: the median 4-bit release changes 2.6% in English and 6.7% to 13.1% in six lower-resource languages, and no model card documents the calibration seed. Given that prior work has shown that the calibration language impacts perplexity and accuracy, but that target-language calibration does not reliably restore downstream behaviour after pruning, we investigated whether the calibration corpus is a cause of these changes. We performed 4-bit GPTQ quantisation on Qwen3-8B, varying only the corpus, with 8 to 16 seeds per corpus under pre-registered criteria. On Belebele, averaged across three option orders, English calibration reduces accuracy by 4.0 to 5.8 points in six lower-resource languages; in contrast, target-language calibration maintains 1.9 to 3.3 of these points. Spanish calibration does not benefit Basque, confirming that the effect is language-specific. A 25% share of the target language achieves most of the reduction, as the Basque curve predicted in three other languages. At the calibration budget of the most-downloaded GPTQ release (eight times ours), a balanced seven-language corpus keeps 1.5 to 2.2 points across all six languages, English inside a registered margin. The effect holds in Llama-3.1-8B-Instruct and is smaller in Gemma-3-12B-it, which handles these languages better. We release the corpora and recipes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.