Federated Fine-Tuning of Code Language Models: A Comparison of FedAvg, FedProx, and FedAdam UnderLanguage-Heterogeneous Clients
Abstract
Large language models specialized for source code excel at program synthesis, but further domain gains require fine-tuning on codebases that are often proprietary and cannot be pooled centrally. Federated Learning (FL) offers a path to collaboratively fine-tune a shared model without exchanging raw data, yet the comparative behavior of FL aggregation algorithms on code generation—particularly under realistic, language-based client heterogeneity—has not been measured systematically. We present a controlled empirical study that fine-tunes Qwen2.5-Coder-1.5B with Quantized Low-Rank Adaptation (QLoRA) inside the Flower framework, comparing three aggregation algorithms (FedAvg, FedProx, FedAdam) across three data-partitioning strategies (random IID, Dirichlet α=0.5, and a realistic non-IID split where each client specializes in one programming language). The resulting global models are benchmarked against a centralized upper bound on HumanEval+, MBPP+, and MultiPL-E using greedy pass@1, with three seeds per condition and paired significance testing. Two findings emerge. First, on algorithm choice, FedAvg consistently matches or exceeds FedProx and substantially outperforms FedAdam in every partition (HumanEval+ pass@1 of 40.0, 36.4, 29.1 for FedAvg across IID, Dirichlet, and by-language), whereas FedAdam is unstable and collapses to 11.6±5.8 under quantity skew; contrary to the expectation that server-adaptive optimization best handles heterogeneity, the simplest aggregator is the most accurate and robust. Second, on the cost of federation, the best federated model is statistically indistinguishable from the centralized reference (38.2) under IID and Dirichlet but lags by ≈9 points under the by-language split. A per-language MultiPL-E analysis reveals a stable competence gradient from Python and JavaScript down to Go, without collapse toward a dominant language. We release adapters and configurations for reproducibility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.