Minimal-State Optimization: How Little Optimizer Memory Does AdamW’s Accuracy Need?
Abstract
Adaptive optimizers such as AdamW store two extra tensors per parameter (first and second moments), doubling the model’s memory footprint. How much of this optimizer state is actually necessary to match AdamW’s accuracy? We study this through A∗Grad, a learning-rate-free optimizer whose step size is set by a Polyak rule with a trust re- gion, and whose only tunable hyperparameter is a dimensionless ratio ρ. By varying only the geometry of the update direction, A∗Grad spans a spectrum of optimizer-state sizes. We find the required state is task-dependent, and a sharpness diagnosis predicts it: the sign-based step (zero state) discards gradient magnitude and converges to min- ima 5–10× sharper than SGD’s (filter-normalized perturbation), which hurts generaliza- tion. Restoring magnitude resolves this issue, though the optimal restoration mecha- nism is strictly task-dependent. For from-scratch vision (CIFAR-100/WRN), a global L2- normalized step with weight decay suffices at zero optimizer state, outperforming AdamW by +2.4%p. For Transformer fine-tuning (BERT/MNLI), per-coordinate RMS is needed; at half AdamW’s state, it outperforms it by +0.5%p with the smallest seed variance among all methods. We place A∗Grad against two recent low-memory optimizers, MicroAdam and NanoAdam, on an accuracy–memory Pareto front, and delineate the operational regimes where A∗Grad demonstrates clear advantages (stability, zero-state vision, hyperparameter- free operation) versus scenarios where its benefits are attenuated (certain fine-tuning tasks, peak accuracy at scale). The code of the proposed optimizer is publicly available at Github (https://anonymous.4open.science/r/astar_grad-D001).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.