acceptodds
Under review as a conference paper at ICLR 2027

Minimizing Label Distribution Shift through Client Selection in Federated Learning

Abstract

Federated learning (FL) is a distributed learning approach that trains a machine learning model using data from multiple clients without exposing each client's local data. Although standard FL assumes a uniform test label distribution, many practical applications, such as medical diagnosis, involve known or estimated test label distributions that often deviate from uniformity. Existing client selection strategies are either limited by client's label distribution estimation or restricted to uniform test label distributions, and furthermore, they lack rigorous theoretical guarantees in non-convex settings. To bridge this gap, we propose FedMLS, a novel client selection strategy that minimizes the Kullback-Leibler (KL) divergence between the test label distribution and the time-averaged label distribution of the selected clients. We establish convergence guarantees for FedMLS under test label distribution shift and extend our analysis to scenarios with loss reweighting. Moreover, we design an iterative mechanism leveraging Fully Homomorphic Encryption (FHE) to implement FedMLS without exposing sensitive clients' label statistics. Extensive experiments on the CIFAR-100, iNaturalist-User-120k, and Shakespeare datasets demonstrate that FedMLS consistently outperforms state-of-the-art baselines, achieving higher test accuracy and up to 10.7× faster convergence than FedAvg (4.6× when combined with FedIR loss reweighting).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.