acceptodds
Under review as a conference paper at ICLR 2027

Layer-wise Adaptive Gradient Compression with Generalized Error Estimation

Abstract

Reducing gradient communication is critical for large-scale data-parallel training. While layer-wise adaptive compression schemes can outperform uniform layer-wise ones, they typically rely on distance in gradient space as the error measure with a fixed measurement schedule. We propose Generalized Gradient Compression (GenComp), which generalizes layer-wise adaptive compression along three axes: 1) the space in which compression error is measured, on gradients or on latent activations obtained by simulating gradient updates and forward passes; 2) the error measure; 3) the frequency of measurement. We show that gradient- and activation-based error measurements, in tandem with error measures addressing information loss and different frequencies of measurement, induce distinct layer-wise compression dynamics, and can improve task performance and compression ratio. Across sparsification, quantization, and low-rank compression schemes applied to different model architectures, GenComp achieves significantly higher compression ratios than recent layer-wise adaptive schemes while retaining similar performance, which leads to reduced end-to-end training time for larger models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.