acceptodds
Under review as a conference paper at ICLR 2027

Ano: Adaptive Optimization for Tracking Under Noise and Drift

Abstract

Deep reinforcement learning asks optimizers to track moving objectives from persistently noisy gradients. We show that, in this regime, Adam's momentum improves the reliability of its update direction but not its progress. Because Adam's update is linear in the momentum, the variance reduction it buys is spent on shorter steps rather than on progress along the signal: at a fixed learning rate, Adam's signal-aligned progress per step matches that of RMSprop, which uses no momentum at all. This trade-off is sound when the objective is fixed and more steps can compensate, but costly when the objective moves. We introduce Ano, which uses momentum only to choose the direction of each coordinate, not the size of the step: the update takes the sign of the momentum and the magnitude of the current normalized gradient. The variance reduction then goes into direction rather than into shorter steps, roughly tripling signal-aligned progress per step in the high-noise limit. Applying the same logic to the second-moment estimate yields an Active Contraction Accumulator, which returns to the post-burst gradient scale in finite time where Adam's estimate returns only asymptotically. Ano keeps Adam's state and hyperparameter set. On real SAC gradient streams, it produces larger signal-aligned updates than Adam. Across 10 MuJoCo and Atari environments with 25 seeds, it improves aggregate return over Adam and four other adaptive and sign-based optimizers, while remaining competitive rather than superior on stationary supervised benchmarks, where shorter steps act as useful annealing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.