TetraBack: Identifying and Correcting Optimization Bias in Microscaled 4-bit Training
Abstract
Microscaled 4-bit weight and activation quantization (W4A4) accelerates LLM inference, but maintaining accuracy at such low precision remains challenging. Activation quantization is a major source of this degradation, yet its effects on training dynamics remain underexplored. In this work, we show that this accuracy gap stems partly from optimization bias in the backward approximation, rather than solely from the limitations of low-bit representation. We consistently observe sustained growth in activation magnitudes and pre-FFN RMSNorm weights during standard MXFP4 W4A4 quantization-aware training (QAT), a pattern that impairs optimization and is absent when activations remain unquantized. We trace this behavior to the conventional straight-through estimator (STE) used for activation quantization, which ignores the gradient contribution from each group's scale and biases gradients reaching the preceding RMSNorm. We propose **TetraBack** to restore this missing gradient contribution without modifying forward computation. Our method curbs the abnormal growth of activation magnitudes and RMSNorm weights, consistently improves performance across 3B–94B MoE models, and reduces the loss gap induced by activation quantization by up to **41%**.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.