When Do Stale Updates Help? Forecasting, Participation Weighting, and Step Scale in Federated Learning
Abstract
Stale client updates can reduce the variance of federated aggregation, but accuracy gains do not by themselves identify the value of forecasting: participation weighting and aggregate scale can change simultaneously. We analyze these effects through a forecast–residual estimator and use OPTICAST, which learns client-specific linear forecasts, as a case study. Exact conditional moments expose the coupling between propensity error and forecast residuals. A complementary decomposition separates aggregate weight mass from relative client weighting, and an aggregation-aware damping criterion identifies the optimal scalar correction at a frozen history. Across the four main image-classification settings, OPTICAST achieves higher mean accuracy than FedAU, MIFA, U-FedVARP, and FedStale. The control experiments qualify this result. Across three CIFAR-10 participation regimes, last-update reuse yields paired gains of 0.14, 1.06, and 3.51 percentage points over zero forecasting as participation decreases, whereas a separate five-seed test does not establish an advantage of learned gains over simple reuse. On FEMNIST, a one-seed oracle scale intervention recovers 14.34 points of a 16.54-point zero-control-to-participating-mean gap. Learned and damped configurations also collapse on Shakespeare. These findings distinguish configured-method performance from evidence for predictor complexity, and motivate evaluating forecasting jointly with participation weights, aggregate scale, implementation consistency, and stability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.