Alternating LoRA for Decentralized Fine-Tuning: A Misalignment–Staleness Trade-Off
Abstract
Low-rank adaptation (LoRA) offers a communication-efficient approach to decentralized federated fine-tuning, but mixing independently trained LoRA factors can create discrepancies in their effective updates. Alternating factor updates mitigate a related aggregation problem in centralized federated learning, where server synchronization maintains a shared frozen factor; neighbor-only mixing does not provide this guarantee. We propose TAD-LoRA, which alternates local updates of the two factors in fixed-length phases while jointly mixing both factors at each communication round. Joint mixing reduces frozen-factor disagreement with the same per-round communication payload as standard decentralized LoRA. Under an idealized one-step stochastic-gradient model, we derive a finite-horizon bound for the loss of an averaged effective adapter. The bound separates a phase-averaged factor-product discrepancy that decreases with the switching interval from a delayed block-coverage term that increases with it; the minimizer of a loose upper-bound surrogate depends on network mixing. On four GLUE tasks, a matched-\(T=1\) ablation shows that joint mixing generally improves over active-factor-only mixing, while inter-client similarity measurements support improved factor alignment. Fixed-interval sweeps reveal that favorable switching regimes depend on network connectivity, with the largest best-observed gains over decentralized baselines under sparse communication. Additional experiments examine the method across multiple data-heterogeneity levels, communication topologies, and client counts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.