acceptodds
Under review as a conference paper at ICLR 2027

FLAME: First-Moment-Free Low-Rank Gradient Projection with Adam-Style Convergence Guarantees

Abstract

Low-rank gradient projection has emerged as a promising approach to optimizer-state compression, but existing methods still suffer from unstable optimization and high computational overhead. To understand these limitations, we develop a unified analytical framework and prove that, under independent random projections, first-order momentum retains only within-period history whereas the second-order state contains information from all periods, and that the induced adaptive updates effectively preserve direction in expectation. Guided by these findings, we propose FLAME (First-Moment-Free Low-Rank Adaptive Moment Estimation), which updates parameters using the current Gaussian-projected gradient normalized by the previous low-rank second-order estimate. It enables effective optimization while halving optimizer-state storage relative to existing Adam-based low-rank gradient projection methods and substantially reducing computational overhead. We prove an convergence rate for FLAME under discrete-time Adam-style optimization. FLAME further incorporates activation compression to reduce training memory. Experiments on LLaMA-3 fine-tuning and pretraining demonstrate competitive task performance and favorable convergence behavior. In terms of system efficiency, FLAME achieves the lowest peak memory among competing low-rank gradient-projection methods, reducing memory by 32-53% relative to full-rank Adam, while improving throughput by 23-45% over the fastest competitor.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.