Q-MINO: A Minimal-Norm Method for Quantization-Aware Training
Abstract
The Straight-Through Estimator (STE) is a widely used heuristic for Quantization-Aware Training (QAT), but its surrogate gradients can exhibit substantial mismatch with the underlying quantized objective, leading to noisy updates and parameter oscillations, particularly in ultra-low-bit regimes. We propose the Quantization-Aware Minimal-Norm Optimizer (Q-MINO), a temporal bundle method that combines gradient consensus, state-drift regularization, and an alignment constraint to construct stabilized, minimum-norm update directions from recent optimization states. Q-MINO solves the resulting constrained subproblem using a warm-started Frank–Wolfe procedure with a feasible fallback initialization. Theoretically, via a stochastic quasi-Lyapunov Kurdyka–\L ojasiewicz (KL) framework, we show that Q-MINO achieves almost sure convergence to a stationary point. Moreover, we detail numerical experiments with Q-MINO at various quantizations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.