TTA-GT: Time-to-Accuracy Oriented Adaptive First/Second-Order Gradient Tracking for Decentralized Training
Abstract
Decentralized training of deep neural networks has emerged as a promising paradigm for scalable AI development, offering the potential to overcome the scalability limitations of conventional centralized training infrastructures. Decentralized training efficiency is often measured in terms of iteration count. This metric can mis-represent wall-clock efficiency when first-order (FO) and second-order (SO) directions have different runtime costs. We study time-to-accuracy (TTA) oriented direction selection in decentralized gradient tracking (GT). Each node chooses between a low-cost FO direction and a damped SO direction. We propose TTA-GT, a node level selector that estimates a local TTA score for each candidate. The score is the product of predicted remaining iterations and per round time, which decides whether the expected future iteration saving of SO justifies its current runtime cost. Prediction affects efficiency but not correctness, with convergence guaranteed by explicit direction constraints. Under standard strongly convex and smooth assumptions, we prove linear convergence for adaptive FO/SO direction sequences whose applied directions satisfy explicit alignment and norm conditions. Experiments on synthetic quadratic problems and logistic regression on real datasets show that TTA-GT can reduce end-to-end TTA under heterogeneous environments, where nodes vary communication conditions, such as bandwidth and latency, as well as different computation capabilities. In a 300-dimensional quadratic setting, TTA-GT reduces TTA from 2.714 to 1.934, reaching a 28.7% improvement. TTA-GT can achieve significant TTA reduction when SO directions offer sufficient iteration savings, while their additional computation cost remains non-negligible.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.