acceptodds
Under review as a conference paper at ICLR 2027

Understanding and Accelerating the Convergence of Federated LoRA

Abstract

Federated fine-tuning of large pretrained models often requires substantial communication and computation, making parameter-efficient methods such as LoRA attractive. A natural approach is FedAvg(LoRA), which combines FedAvg with LoRA. Despite its practical utility, the convergence theory of FedAvg(LoRA) remains limited: existing analyses typically assume bounded LoRA factors to ensure the smoothness of the factorized objective, an additional assumption not guaranteed by LoRA updates. More recent work on centralized LoRA shows that LoRA gradient descent achieves complexity without this assumption. This raises the question of whether FedAvg(LoRA) can similarly achieve communication complexity without this assumption. Unlike centralized LoRA, FedAvg(LoRA) performs multiple local steps that introduce client drift, and this drift can be amplified as the factor norms grow. We resolve this analytical challenge and establish an communication upper bound under full client participation. We then prove a lower bound for multiple local steps, even when the stepsize is chosen adaptively from the optimization history. To accelerate this communication complexity, we introduce FedStabLoRA, which uses persistent client corrections to reduce client drift and stabilize training via a quadratic penalty. FedStabLoRA achieves communication complexity without a bounded-factor assumption, and the rate also holds under partial participation. Across seven language and vision benchmarks, FedStabLoRA achieves the best performance among the compared methods. Its advantage persists under severe data heterogeneity, more local training, and partial participation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.