acceptodds
Under review as a conference paper at ICLR 2027

When Is Optimizer Memory Federable? Adaptive Synchronization in Federated Learning

Abstract

Federated learning methods increasingly share optimizer states across clients, yet these states are synchronized on the model's clock or on a schedule fixed by the method. We ask which optimizer memories should be federated, and how often. Decoupling optimizer-memory synchronization from model synchronization, we distinguish federability, the immediate benefit of replacing a client's memory with the aggregate, from the long-run utility of a synchronization policy. Fixed-interval experiments show that Adam's second-moment tolerates long intervals, whereas the FedMuon policy we evaluate degrades once the interval is expanded under strong heterogeneity, largely through the coupling of momentum delivery with its global-direction correction. For Adam, an exact first-step analysis and measurements on training runs show that the aggregate effect of transporting the second-moment depends on the joint variation of client memory deviations and optimizer response rather than on memory discrepancy alone, and that the realized gain follows the signed alignment between the endpoint difference and the evaluation-loss gradient. We therefore introduce a matched transport-gain probe that compares short-horizon training with local versus aggregated memory, and adapt each memory's interval from its outcome. On CIFAR-10 and CIFAR-100, the controller recovers the fixed-interval sensitivities and reduces uplink by up to , including probe traffic, while staying within about one point of every-round synchronization. Code is available at the anonymous repository.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.