Alignment Paradigm of Intrinsic Subspace Experts for Federated Fine-Tuning
Abstract
The rapid scaling of foundation models has made full-parameter training increasingly impractical in federated scenarios due to prohibitive computational and communication costs, motivating the development of federated parameter-efficient fine-tuning (PEFT). Meanwhile, mixture-of-experts (MoE) based adaptation has shown promising capability in improving model specialization and scalability. However, directly applying MoE-based LoRA adaptation to federated fine-tuning remains challenging. The fundamental difficulty lies in various forms of misalignment introduced by decentralized optimization. Furthermore, MoE architectures introduce additional challenges under heterogeneous clients, where experts with identical positions may capture inconsistent knowledge and independently optimized routers may become difficult to aggregate. To address these challenges, we propose A-SMoE for federated fine-tuning. Instead of allowing local experts to freely learn independent adaptation spaces, A-SMoE exploits the intrinsic subspaces of pretrained models through singular value decomposition and freezes the corresponding input and output bases. This design provides a shared semantic coordinate system for all clients, thereby improving the consistency of expert representations and enabling more reliable federated aggregation. Moreover, we introduce a parameter-free routing mechanism based on subspace projection energy, eliminating additional router parameters and avoiding auxiliary objectives required for router balancing. Extensive experiments on multiple benchmark datasets demonstrate that A-SMoE consistently outperforms existing federated fine-tuning methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.