SpaRFL: Sparse Low-Rank Federated Learning for Memory-Constrained Clients
Abstract
Federated Learning (FL) enables collaborative training across decentralized data sources, but standard FL methods typically assume that each participating client can store and optimize the full global model. We study federated Transformer training in the no-client-fits regime, where no participating client can store and train the full dense global model within its local memory budget. To address this setting, we propose SpaRFL (Sparse Low-Rank Federated Learning), a client-specific one-shot sparse+low-rank projection method that constructs trainable client models under heterogeneous memory constraints while preserving a dense server model. Across vision, natural language understanding, and language modeling, SpaRFL consistently outperforms the considered model-heterogeneous baselines under severe resource constraints. Under a projected-linear training-state budget 90% below dense training, SpaRFL remains within 1% of the dense baseline on MRPC. In federated pretraining on C4, LLaMA-130M trained at a 70% reduction in the same budget achieves lower perplexity than the dense LLaMA-60M baseline. Overall, SpaRFL provides a practical approach to federated training under strict client memory constraints.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.