FedGPA: What to Adapt and What to Share in Personalized Federated Fine-Tuning
Abstract
Personalized federated fine-tuning allows clients to adapt pre-trained models to local data while benefiting from knowledge shared across clients, but repeatedly synchronizing model updates can incur substantial communication cost. Existing methods typically make adaptation and sharing decisions at the granularity of complete layers, adapters, or other coarse components, overlooking the fact that different attention heads and feed-forward network (FFN) units within the same Transformer layer can contribute differently to downstream adaptation. We propose FedGPA, a personalized federated fine-tuning framework that jointly determines adaptation placement and sharing roles at the granularity of intra-layer Transformer units. FedGPA constructs a unified, cost-aligned candidate space of attention heads and FFN clusters, estimates the benefit of each candidate for global sharing and local adaptation, and assigns each unit to a Global, Private, or Frozen role. To limit the overhead of fine-grained selection, FedGPA performs a one-time compressed role allocation; during subsequent federated training, only the always-shared task head and Global units are synchronized, while Private units are adapted and retained locally. Experiments with BERT-base and RoBERTa-base under non-IID federated settings show that FedGPA maintains strong final test performance while substantially improving communication efficiency, reducing the cumulative communication required to reach a target accuracy by up to 76.5% compared with the personalized FedDPA baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.