Beyond Zero-Padding: Nested-SVD Aggregation for Heterogeneous-Rank Federated LoRA
Abstract
Federated fine-tuning with LoRA typically assumes all clients share one adapter rank, even though heterogeneous devices may need different ranks to meet local memory and compute constraints. The common remedy, zero-padding smaller LoRA factors to a shared maximum rank, preserves a uniform parameter shape but communicates the update in a larger shape than it needs. We propose NSVD-Agg, a nested-SVD aggregation method that instead projects each client’s native-rank update onto a shared reference subspace of dimension \(r_\max\). The subspace is estimated from the raw client updates and refreshed across rounds, while the current basis stays fixed during aggregation – a temporal ordering required for the non-convex convergence guarantee we prove. On Llama-3.2-3B with simulated heterogeneous ranks \({4,8,16,32}\), NSVD-Agg achieves lower final loss than zero-padding in all seven random seeds while reducing client-to-server LoRA communication by an analytical \(53.1%\); the comparison replicates on Qwen2.5-3B, where the improvement is larger and more consistent. We further introduce an untracked exact-product baseline that shares NSVD-Agg’s communication savings but, by construction, lacks its guarantee; across seven seeds on both models it matches or slightly exceeds NSVD-Agg’s loss, indicating the gain over zero-padding comes primarily from exact-product aggregation rather than from tracking itself. Tracking’s role is therefore to supply a proven convergence guarantee at a small, measured empirical cost. NSVD-Agg is thus a communication-efficient, theoretically grounded aggregation scheme whose empirical advantage and theoretical guarantee arise from distinct, separately-verified mechanisms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.