acceptodds
Under review as a conference paper at ICLR 2027

Beyond Aggregate Drift: Optimal Client-Specific Windows in Federated Learning

Abstract

We study federated learning from continuously arriving data: each client samples from its own distribution, these distributions drift over time, and the server must track the minimizer of the clients' weighted average loss. Clients are not alike as they collect data at different rates, with different noise, and change at different speeds. The main question is how much history should each one keep? A single shared window must serve every client at once, and so makes a fast-changing client carry the sampling noise of a slow, nearly stationary one. We give each client a memory of its own, sized to that client's noise, arrival rate, and drift, and refresh it by restarted accelerated method that starts from a common point. The tracking error is then bounded by \[ R\left(\sum_m w_m^3/2 \sigma_m\gamma_m/\lambda_m\right)^2/3, \] for clients with weight , noise , arrival rate , and drift rate on a domain of diameter . This is never worse than what any common window can achieve, and strictly better whenever noise and drift are spread unevenly across clients. A lower bound that puts different groups of clients on different drift clocks shows this quantity is the right one. The two bounds agree up to a constant whenever the clients occupy boundedly many refresh scales, and in general differ only by the sixth root of the number of scales. Experiments show the gain comes from the memory rather than from the solver: it persists across accelerated, stochastic-gradient, and one-step methods, and against a separately tuned common-memory baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.