Distributed Optimization with Direction-Constrained Biased Compression
Abstract
Error feedback (EF) is widely used to ensure convergence in distributed optimization with biased gradient compression. However, EF requires each node to maintain at least one additional model-sized numerical vector, introducing extra memory overhead. To reduce this memory overhead, we investigate an alternative approach that controls compression errors through directional constraints rather than numerical error accumulation. Our descent analysis identifies a sufficient condition for convergence based on the alignment between the global gradient and the aggregated compression error. Specifically, we introduce ACTION (Acute-Angle CondiTION), which requires their inner product to be non-negative. Since individual nodes cannot access the current global gradient before communication, we further propose L-ACTION (Lagged ACTION), which uses the previous aggregated compressed gradient as a reference. By enforcing the directional constraints coordinatewise, L-ACTION requires retaining only the signs of this reference as auxiliary state for compression error control, substantially reducing the memory footprint compared with the numerical error states maintained by EF. We establish convergence guarantees for ACTION under generalized biased compression without requiring compressor contractivity and show that L-ACTION achieves the same convergence rate. Mainstream compression methods can be modified to satisfy the directional constraints with minor adjustments. Experiments demonstrate lower training loss, higher test accuracy, and improved communication efficiency in several settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.