acceptodds
Under review as a conference paper at ICLR 2027

PowerStep: Memory-Efficient Adaptive Optimization via -Norm Steepest Descent

Abstract

Adaptive optimizers such as Adam are standard for training Transformers, but storing gradient first and second moments incurs substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by -norm steepest descent, PowerStep applies a signed-power transform directly to one momentum buffer. We establish a finite-horizon stationarity bound for exact, unregularized updates, with an term and a noise-dependent residual. Experiments on Transformers from 124M to 235B parameters show competitive validation quality while halving optimizer-state memory. Combined with uniform quantization, PowerStep remains numerically stable and reduces optimizer-state memory by compared to AdamW. PowerStep thus provides a simple, memory-efficient alternative for large-scale training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.