Two-Level Normalization for Federated Optimization under Generalized Smoothness
Abstract
We study federated optimization with two-level normalization: clients perform normalized stochastic local updates, and the server applies a normalized update along the aggregated client displacement. This structure is natural under generalized -smoothness, where gradient magnitudes may be large and classical smoothness-based FedAvg analyses do not directly apply. The main difficulty is that exact normalization introduces a bias: the server direction is no longer an unbiased gradient estimate, and the bias depends on the dispersion of stochastic gradient norms within a communication round. We make this effect explicit through a norm-dispersion parameter . Under standard stochastic-noise and heterogeneity assumptions, we prove two-phase optimization rates up to an explicit error floor with client-and-step averaging of the stochastic term. In the non-convex setting, the rate contains both the standard sublinear component and an improved generalized-smoothness component. In the convex setting, we derive both sublinear and linear optimization behavior, up to the same bias/noise floor. We further extend the guarantees to mini-batching, partial participation, and heavy-tailed local stochastic noise, showing how batching and participation reduce stochastic and sampling-heterogeneity terms while the normalization-bias term remains explicit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.