Adam as Magnitude-Weighted Sign Descent: Dimension-Free Convergence under Weak Assumptions
Abstract
We develop a convergence analysis of fully bias-corrected Adam through its *magnitude-weighted stochastic sign* representation. The key step controls the accumulated wrong-sign loss while retaining the joint randomness of Adam's update magnitude, momentum, and gradient noise. Under -smoothness and conditional -affine second moments, we prove that the expected average gradient norm is at most , where has no dependence on explicit dimension factor and the rate has no extra ; under global point-pair -smoothness, a stopped analysis yields the same exponent with high probability. We further instantiate the theory for fan-in-scaled LayerNorm/RMSNorm networks with fixed depth and increasing hidden width. For a concrete fixed-affine family, explicit width–horizon conditions guarantee that the Adam iterates remain in a width-uniform parameter region throughout training, yielding a convergence constant independent of the total parameter count. To our knowledge, this is the first Adam convergence-rate result establishing dimension-free optimization for an explicit growing family of normalized networks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.