JacQuant: STE‑Free Quantization‑Aware Training via Learned Jacobian Surrogates
Abstract
Quantization-aware training (QAT) usually propagates gradients through hard quantizers with a fixed straight-through estimator (STE). At ultra-low precision, this backward rule ignores how quantized weights respond to finite perturbations. We introduce JacQuant, a QAT framework that learns a lightweight surrogate of quantizer sensitivity while retaining the hard forward quantizer. A periodic Gaussian probe fits one scalar gain per weight group; the gains are averaged over time and reused between refreshes. We identify the population regression target through Gaussian smoothing and separate it from the clipped estimator used in practice. For a local linearized surrogate, we establish non-convex stationarity bounds and linear convergence to an error neighborhood under a Polyak–Łojasiewicz condition, with explicit bias and variance terms. Across 1B–8B LLaMA and Qwen models with 1–2-bit weights, JacQuant improves both perplexity and mean zero-shot accuracy over the ParetoQ and WinQ baselines. On LLaMA3-1B W2A16, it lowers ParetoQ perplexity by 8.9% and recovers roughly one third of the gap to FP16 in both metrics. Gains persist at all three 8B checkpoints through 90K steps. In its default configuration, JacQuant incurs only a 1.6–3.5% reduction in training throughput on 1B models and 0.76% on 8B relative to the corresponding STE baselines. It requires no additional computation at inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.