Hadamizer: Spatio-Temporally Adaptive Hadamard Rotation for Low-Precision Distributed Optimization
Abstract
Large Transformer models have become increasingly vital across numerous applications, yet training them is fundamentally bottlenecked on high-bandwidth memory (HBM) with 32-bit optimizer states occupying a dominant share. Quantization is a key technique to alleviate this pressure, but standard coordinate-wise quantization suffers from severe clipping errors caused by heavy-tailed outliers that widen the quantization range. While orthogonal rotations effectively disperse outliers, mixing coordinates across distributed systems makes collective communication prohibitively expensive. Here, we propose HADAMIZER, an optimizer-native orthogonal rotation framework designed for distributed training on accelerator meshes. Spatially, it aligns rotation blocks with local shard boundaries to avoid rotation-induced collectives. Temporally, HADAMIZER adapts the randomization of Hadamard transformations to tensor dynamics: target-adaptive keying employs static keys for persistent optimizer states and dynamic keys for stochastic gradients, preserving preconditioning with minimal quantization errors. Evaluated on TPUv5e meshes, HADAMIZER achieves near-full-precision-equivalent results on various pre-training and fine-tuning tasks, while delivering up to 9.58% higher end-to-end training throughput over the current rotation-based training frameworks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.