Understanding Top- Sparsification in Distributed Deep Learning with LaplaceK
Abstract
Top- sparsification reduces the communication cost of distributed stochastic gradient descent, but existing convergence analyses often characterize its error using the looser Random- contraction. We analyze error-compensated gradients and show that their coordinate magnitudes exhibit a Laplace-like structure, which leads to a tighter characterization of the error introduced by Top- selection. Based on this observation, we introduce LaplaceK, a GPU-oriented selector that uses a Laplace-tail prior to localize a small candidate pool, performs exact Top- selection within that pool, and falls back to global selection when the candidate pool is invalid. This design preserves the exact selected coordinates and transmitted values while reducing selection overhead. Experiments across six non-Transformer image-model configurations show that LaplaceK maintains convergence and final accuracy comparable to magnitude-based sparsification while reducing wall-clock training time. Across the reported experiments, LaplaceK achieves up to speedup over standard TopK-SGD.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.