acceptodds
Under review as a conference paper at ICLR 2027

Data-Adaptive Transforms Can Improve the Accuracy of Quantization-Aware Training

Abstract

Performing large model training in lower precision, a process known as quantization-aware training (QAT), has become a standard approach for reducing costs, both for inference (since it produces a smaller model) and training time (since the matrix multiplications can be performed faster). However, training models accurately in lower precision is still a challenge. In this work, we propose a new, more accurate approach for QAT, based on the idea of adaptively learning a set of “smoothing” matrices for each layer's matrix multiplication, whose goal is to specifically reduce its quantization error. We show for the first time that this adaptation can be performed efficiently and stably during training, leading to significant reductions in the accuracy gap for QAT for pre-training, fine-tuning, and distillation of quantized large language models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.