GIFT: Geometry-Informed Low-Precision Gradient Communication for LLM Pretraining
Abstract
Gradient communication is a major scaling bottleneck in large language model (LLM) pretraining. Low-precision formats such as FP8 can substantially reduce communication cost, but often at the expense of gradient fidelity. Existing approaches either quantize gradients directly in their original coordinates or apply data-independent transformations to suppress outliers. Our study shows that both approaches overlook a fundamental source of quantization distortion: LLM gradients are highly anisotropic, with geometry that varies substantially across layers, making low-precision error strongly direction dependent. This observation suggests that gradient fidelity depends not only on the quantizer itself, but also on the coordinates in which quantization is performed. Motivated by this insight, we design GIFT, a geometry-informed gradient communication method that transforms gradients into more isotropic coordinates before low-precision quantization. To make this transformation practical, GIFT uses a one-sided, low-rank approximation of the geometry and applies it only to layers that are most sensitive to quantization errors, while leaving the optimizer, training recipe, communication collective, and low-precision format unchanged. Across Llama-style models with approximately 600M and 2B parameters, GIFT achieves downstream task performance comparable to Muon training with FP32 gradients. On 64 NVIDIA GH200 Superchips, GIFT reduces end-to-end training step time relative to FP32 gradient communication by 10.87% for Llama-600M and 17.05% for Llama-2B.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.