Who Pays for Compression? The Answer Depends on the Instrument
Abstract
Post-training quantization (PTQ) degrades some languages far more than others, and the standard explanation blames the calibration set: calibration-based methods such as GPTQ are said to push quantization error out of the subspace the calibration data spans, so languages far from that subspace should suffer most. We show that the mechanism is real, that it is protective rather than harmful, and that the evidence for the attribution is far more fragile than the literature assumes. We confirm the mechanism directly: GPTQ drives the normalized in-subspace residual energy γ̄ to 0.213–0.277 while round-to-nearest and norm-matched isotropic noise sit at 1.00, across 3 multilingual encoders and an autoregressive decoder. Evacuation is beneficial: at identical weight-error norm, GPTQ does 1.51×–11.61× less damage than isotropic noise. We then ask whether a pre-computable misalignment statistic R_ell predicts which languages pay, and whether that prediction is specific to calibration-based quantizers. Using quantized weights that are bit-identical across conditions, we show the answer depends on the measurement instrument. Under non-parallel evaluation the apparent calibration-specificity is inflated; under content-controlled parallel evaluation it disappears on encoders scored by a masked-language-model proxy, yet remains large on a decoder scored by bits-per-byte. We argue that per-language quantization damage must be measured on parallel corpora with the deployment objective, and we quantify how many languages such claims actually require.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.