acceptodds
Under review as a conference paper at ICLR 2027

Where Low Precision Bends Gradients

Abstract

Low-precision training reduces training time and memory use for transformers. Realizing these gains requires deciding which tensors, operations, and matrix-multiplication passes retain higher precision. Existing precision-allocation methods have focused on local reconstruction error or forward-loss sensitivity. These criteria measure changes in tensor values or the forward loss, but do not directly measure the resulting change in the parameter gradient that drives optimization. We reveal this mismatch by measuring how quantization errors in forward and backward computation change the full model's parameter gradient: Similar local errors at different tensor locations produce gradient shifts whose magnitudes differ by more than tenfold. These shifts mainly rotate, or bend, the gradient. To preserve gradient direction during low-precision training, we introduce bending-guided precision allocation. The method allocates high precision to selected operation groups or individual matrix-multiplication passes where quantization rotates the gradient. Using only a small amount of calibration data, the proposed method determines how much high precision is needed and where to allocate high precision to bring the gradient directions of 4- and 8-bit training close to the bfloat16 (BF16) reference. In GPT-2 continued pretraining, our calibration probe ranks precision allocations before training, closely matching the allocations' final-loss ranking with a Spearman correlation of 0.95, compared with 0.35 for local reconstruction error. Under BF16 mixed-precision evaluation, a model trained in 8-bit floating point (FP8) with our allocation matches the loss of a BF16-trained model. These allocations use 11% to 24% less high-precision activation-cache memory than a fixed policy that stores the output activations of the token embedding, output head, and every normalization layer in high precision. In matrix-multiplication simulations of NVIDIA's NVFP4 format, our probe identifies forward passes as the main source of bending. Executing forward multiplications in BF16 uses high precision for just one third of the transformer-block matrix-multiplication work and closes 99% of the loss gap to BF16 training. At a matched high-precision compute budget, allocating BF16 to selected forward matrix multiplications closes 52% of the gap, compared with at most 20% for NVIDIA-style block protection. The method also improves low-precision training across four vision backbones.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.