LACCO: Local Accumulation with Computation–Communication Overlap
Abstract
Training language models across low-bandwidth networks requires reducing synchronization overhead. Local optimization methods such as DiLoCo communicate only after several local steps, but outer gradient averaging and the outer update remain on the critical path. Existing computation–communication overlap methods avoid this stall through delayed or replica-local updates, which can degrade convergence or cause replicas to drift. We propose LACCO (ocal ccumulation with omputation–ommunication verlap), a two-stage algorithm that overlaps global communication with local computation while applying globally aggregated updates shared by every replica, without introducing staleness in the outer update. We establish convergence guarantees for smooth non-convex objectives. Experiments on language-model pre-training with FineWeb show that LACCO consistently outperforms communication-overlapped baselines such as Eager and Delayed DiLoCo. These gains remain robust as model size scales up to 325M parameters and across a range of communication intervals and numbers of replicas up to 16 workers, while incurring only negligible—or fully overlappable—computational or memory overhead, which demonstrates the strength of LACCO.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.