acceptodds
Under review as a conference paper at ICLR 2027

Cumulative-Variance-Adaptive Decentralized STORM

Abstract

Gradient noise in decentralized optimization may vary along the optimization trajectory, while the smoothness constant is often unknown. We propose Noise-Adaptive Decentralized STORM, which combines locally accumulated direction energy, a two-sample noise proxy, and a neighbor max-envelope to jointly adapt stepsizes and gradient-estimator refresh coefficients. Its horizon-independent fixed hyperparameters require no prior knowledge of the smoothness constant or noise level. Under mean-square smoothness, bounded centered noise, and the stated synchronous-network assumptions, we establish an expected gradient-norm guarantee governed by cumulative expected conditional variance along the optimization trajectory. A uniform variance bound yields a rate of , improving to under bounded cumulative variance, including zero noise, with the same hyperparameters. Experiments show robustness to increased stepsize multipliers and fewer queries to prespecified stationarity thresholds than the evaluated baselines in the tested vanishing-noise environments. Classification experiments additionally characterize the accuracy–resource tradeoff.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.