STE Is Not a Heuristic: A Variational Justification in Quantization-Aware Training
Abstract
The straight-through estimator (STE) is widely used in quantization-aware training (QAT) and other discrete optimization problems in deep learning, but it is considered to be a convenient heuristic with limited theoretical justification. We show that STE admits a principled variational derivation: starting from a variational relaxation of the combinatorial quantization objective under stochastic rounding, STE emerges as the first-order Taylor approximation of the multivariate expansion around the MAP configuration, placing it on rigorous theoretical footing. This variational perspective naturally extends beyond uniform grids: replacing stochastic rounding with a Laplace categorical distribution that is annealed toward the hard solution concentrates gradient signal on uncertain assignments and enables joint learning of non-uniform codebooks whose grid points migrate to minimise reconstruction error, an extension we call non-uniform latent QAT (NUL-QAT). Evaluating on the Qwen3 family (0.6B–32B) at 3-, 2-, and 1.58-bit precision with block-wise QAT, we find that NUL-QAT outperforms STE-QAT on WikiText perplexity and on MMLU, GSM8K, WinoGrande, HellaSwag, and ARC-C accuracy, with the largest gains at extreme bit-widths, confirming that the variational framework captures weight-distribution structure that uniform grids cannot. A custom Triton kernel evaluates the Laplace expectation without materialising the K-fold probability tensor, and the learned non-uniform grids allow efficient inference with lookup table-based kernels such as FLUTE.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.