Quantization Everywhere with Rotated Fourier Surrogate
Abstract
Quantization is essential for efficient neural network deployment, whereas its optimization overwhelmingly relies on the Straight-Through Estimator (STE). However, STE discards the structure of the quantization operator by replacing its derivative with an identity mapping, resulting in a mismatch between the forward and backward passes that can fundamentally limit the achievable quality of quantized models. To address this limitation, we introduce Rotated Fourier Surrogate (RFS), a unified framework that extends STE with a principled gradient estimator. RFS analytically reflects the structure of quantization while remaining stable and computationally efficient. We further provide rigorous theoretical guarantees for RFS, demonstrating well-conditioned gradients and favorable optimization properties. Moreover, RFS is plug-and-play and incurs negligible computational overhead with our optimized kernel. We extensively evaluate RFS across diverse quantization paradigms, spanning integer and floating-point quantization, weight and activation quantization, as well as quantization-aware training (QAT), quantization-aware distillation (QAD), and post-training quantization (PTQ). Our experiments cover a broad spectrum of architectures, including large language models, diffusion language models, vision transformers, and convolutional neural networks, as well as diverse applications including generative recommendation, image classification, pre-training, and reasoning post-training. Across these diverse settings, RFS consistently improves upon strong state-of-the-art quantization baselines, demonstrating its broad applicability across quantization paradigms, model architectures, and applications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.