Not All Blocks Need Full Momentum: Sharpness-Aware Memory-Efficient Optimization for Transformers
Abstract
Momentum is a key component of modern optimizers, but maintaining full-rank momentum for every parameter is increasingly costly for large language model training. We show that momentum is not uniformly valuable across Transformer blocks: different blocks exhibit substantially different local optimization sensitivities, and consequently require different levels of momentum capacity. Motivated by this observation, we propose AdaPM (Adaptive Partial Momentum), a sharpness-aware framework that allocates momentum non-uniformly across parameter blocks. Rather than treating momentum as an all-or-nothing resource, AdaPM assigns each block to one of three regimes—no momentum, low-rank momentum, or full momentum—according to its local sharpness and empirical momentum dependency. For blocks where the dominant optimization dynamics are low-dimensional, we further introduce a residual-corrected low-rank momentum estimator that compensates for information lost by rank truncation while retaining a small memory footprint. This yields a simple principle for optimizer-state reduction: match momentum capacity to the local sensitivity and optimization structure of each block, rather than applying full momentum uniformly. Across GPT-2 and LLaMA pretraining and fine-tuning tasks, AdaPM reduces momentum memory by over 90% while maintaining competitive training performance. When combined with memory-efficient second-order statistics, it reduces total optimizer-state memory by up to 95% and substantially improves training efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.