Intermittent Error Recycling: Narrowing the Gap Between Compressed and Uncompressed SGD
Abstract
Gradient compression is a standard remedy for communication overhead in distributed optimization. Yet, existing unbiased compression methods can substantially degrade convergence by inflating the variance of stochastic updates, creating a gap between compressed and uncompressed training. In this work, we show that this degradation is not fundamental and can be mitigated by exploiting a simple distinction: unlike stochastic gradient noise, compression error is directly observable. We propose Intermittent Error Recycling (IER), a framework in which workers accumulate compression residuals locally and periodically communicate an uncompressed correction. This allows compression errors to be amortized over time rather than repeatedly injected into the optimization dynamics. We show that IER recovers the dominant stochastic convergence term of uncompressed distributed SGD without compression-induced variance inflation, while confining compression-dependent effects to faster-decaying terms. This mechanism does not increase the total communication cost, with the correction frequency determined directly from the communication budget, allowing IER to retain the benefits of compression without its associated variance penalty. We empirically corroborate our theoretical guarantees and demonstrate the efficacy of our approach.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.