How Communication Shapes Convergence and Generalization in Decentralized Training: A Feature Learning Perspective
Abstract
As data, model, and client populations continue to grow, decentralized learning has emerged as an important paradigm for training without a central server. Despite its practical relevance, existing theory has largely focused on optimization and consensus, with generalization typically studied through algorithmic stability. It remains unclear how decentralized communication shapes feature learning in neural networks and, in turn, affects the generalization of individual client models. To address this gap, we analyze decentralized gradient descent for two-layer convolutional neural networks from a feature learning perspective. The derived dynamics of signal and noise coefficients characterize how communication propagates predictive signals and sample-specific noise across clients. The corresponding topology-aware quantities yield early generalization bounds and reveal client-placement effects that graph spectra and feature similarity alone cannot capture. The analysis further establishes training convergence and client-level risk guarantees at a common fitted iterate. An extension to periodic local updates identifies a sufficient local-update budget under which the early generalization guarantee retains the well-mixed scale. Experiments on synthetic and real-world datasets under IID and non-IID settings support the predicted effects of communication topology, client placement, and communication frequency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.