acceptodds
Under review as a conference paper at ICLR 2027

Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization

Abstract

Momentum-based optimizers are widely used in modern deep learning, yet the relations among momentum, update geometry, and acceleration remain only partially understood. We develop an DMM-nspired omentum (AIM) framework based on residual-penalty variable splitting, which interprets momentum as a multiplier-like correction driven by the splitting residual. AIM recovers the exponential moving average of gradients from an ADMM-style multiplier update and separates three mechanisms that are often intertwined in practical optimizers: the residual penalty determines the update geometry, the multiplier update determines how gradient information is accumulated into momentum, and the approximation of the objective-related -subproblem determines the acceleration form. Building on AIM, we propose elativistic daptive gradient escent with ccelerated esidual (RADAR), which combines relativistic adaptive geometry, decoupled residual correction, and gradient-difference momentum filtering to improve the update direction and momentum estimation. We establish stochastic convergence through a variance-perturbed Lyapunov drift analysis. Experiments on supervised vision learning, language modeling, and reinforcement learning show that RADAR achieves consistent improvements over strong adaptive optimizer baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.