QuantRet: Quantization-Robust Language Model Pretraining via Geometric Retractions
Abstract
Low-bit post-training quantization (PTQ) can fail when a small number of coordinates set the quantization range or when high-gain directions amplify local errors. We introduce QuantRet, a pretraining method for residual Transformer decoders. QuantRet retracts selected activations to fixed-radius spheres and constrains the norms of hidden weight matrices throughout pretraining. These operations fix the radial scale at selected activation sites and constrain how hidden linear maps rescale their inputs without changing the quantizer. Our analysis separates weight and activation perturbations, bounds their local amplification under the matrix constraints, and gives a conditional bound on their propagation through spherical residual updates. On Qwen3-like models, QuantRet lowers language-modeling loss under low-bit weight, activation, and key–value-cache quantization, and its strongest variant achieves higher zero-shot commonsense accuracy under the same quantized graphs. Mechanism measurements show that similar weight signal-to-noise ratios can produce different final dot-product signal-to-noise ratios and end-to-end degradation. These results support pretraining-time geometric constraints as a practical way to improve low-bit PTQ.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.