acceptodds
Under review as a conference paper at ICLR 2027

Beyond Exact Equivalence: Rethinking Quantization Preprocessing

Abstract

Most quantization preprocessing methods search for transformations that exactly preserve the pretrained model’s full-precision function. However, the model that matters is the one after quantization: why should optimization remain confined to equivalence before rounding? We challenge this restriction and formulate quantization preprocessing as optimization over a broader space of representations, where exact equivalence is only one possible constraint. Our framework accommodates flexible offline preprocessing, nonlinear parameter transformations, and direct weight adaptation while keeping deployment constraints explicit. We instantiate this perspective with an equivalence-preserving structured initialization followed by global quantization-aware distillation, allowing the full-precision parameters to deviate from the original model to better preserve its behavior after quantization. Across multiple model families, this relaxation consistently enhances post-quantization language-modeling quality compared to matched post-training quantization baselines, without requiring preprocessing-specific online transformations at deployment. These results establish a broader principle for low-bit compression: exact equivalence can provide a safe starting point, but should not define the space in which the quantized model is optimized.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.