Distance-Aware Adaptive Scaling for Stochastic Optimization
Abstract
The ideal learning rate for each model parameter depends on both the magnitude of its stochastic gradients and the distance from its initial value to an optimal value. Given comparable gradient histories, a parameter that starts farther from its optimal value should take larger steps. However, standard adaptive methods such as AdaGrad and Adam adapt to stochastic gradient magnitudes without explicitly accounting for these coordinatewise distances. To address this limitation, we introduce bf distance-aware adaptive scaling, which uses the largest observed parameter displacement from a reference point as a surrogate for unavailable distance information. Coordinates with larger displacement receive larger relative updates, while those with small displacement retain approximately the original scaling. We integrate this mechanism into AdaGrad and AdamW, yielding DA-AdaGrad and DA-AdamW, both of which recover their base optimizers when distance-aware scaling is disabled. For DA-AdaGrad, we derive an AdaGrad-style regret bound and identify conditions under which distance-aware scaling improves the regret-bound expression on a fixed trajectory. Experiments on synthetic and real-world datasets demonstrate the effectiveness of the proposed approach.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.