Convergence of Split Federated Learning via a Masked Compositional Optimization Lens
Abstract
Split federated learning (SFL) trains a split neural network across distributed clients by keeping initial layers on-device and offloading the remaining layers to a server. The split point is called the cut layer, where clients transmit intermediate activations to the server and receive backpropagated gradients. While SFL reduces client compute and memory, the resulting client–server coupling makes its optimization behavior poorly understood. This paper develops, to our knowledge, the first compositional convergence analysis of SFL. We show that SFL optimizes a single-expectation compositional objective, i.e., a client-side forward map produces cut-layer representations, and a server-side map completes the network and evaluates the loss. Unlike existing federated compositional optimization where the main challenge is tracking an inner expectation, SFL's inner map is a per-minibatch network transform, so the key challenge is controlling the split-induced coupling. We address the challenge and establish nonconvex convergence results for two major SFL variants. For the parallel variant SFL-V1, after communication rounds, we obtain an rate. For the sequential variant SFL-V2, we obtain an rate with a leading constant that depends on and split-dependent quantities, where is the number of clients. Enabled by the compositional formulation, our theory reveals insights on SFL design choices, including (i) how the cut layer affects convergence, and (ii) how SFL-V1 and SFL-V2 scale differently with . Experiments on regression and classification tasks corroborate our theory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.