acceptodds
Under review as a conference paper at ICLR 2027

TetraBack: Identifying and Correcting Optimization Bias in Microscaled 4-bit Training

Abstract

Microscaled 4-bit weight and activation quantization (W4A4) accelerates LLM inference, but maintaining accuracy at such low precision remains challenging. Activation quantization is a major source of this degradation, yet its effects on training dynamics remain underexplored. In this work, we show that this accuracy gap stems partly from optimization bias in the backward approximation, rather than solely from the limitations of low-bit representation. We consistently observe sustained growth in activation magnitudes and pre-FFN RMSNorm weights during standard MXFP4 W4A4 quantization-aware training (QAT), a pattern that impairs optimization and is absent when activations remain unquantized. We trace this behavior to the conventional straight-through estimator (STE) used for activation quantization, which ignores the gradient contribution from each group's scale and biases gradients reaching the preceding RMSNorm. We propose **TetraBack** to restore this missing gradient contribution without modifying forward computation. Our method curbs the abnormal growth of activation magnitudes and RMSNorm weights, consistently improves performance across 3B–94B MoE models, and reduces the loss gap induced by activation quantization by up to **41%**.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.