PLUMAGE: A Probabilistic Low-rank Unbiased Min-Variance Gradient Estimator for Memory Efficient Large Model Training
Abstract
Accelerator memory and communication bandwidth are major bottlenecks in large-scale training, motivating low-rank gradient estimators (LGEs) that compress both gradients and optimizer states. Existing LGEs trade bias for variance: deterministic top-k estimators are biased, whereas randomized estimators are unbiased but may incur high variance. We propose PLUMAGE, a probabilistic low-rank unbiased minimum-variance gradient estimator for a given rank budget. PLUMAGE uses fixed-rank sampling without replacement, preserving the asymptotic memory footprint (with only k additional scale factors per layer) of prior one-sided low-rank methods. Additionally, we introduce a numerically robust second-moment alignment procedure for alternating subspaces in stateful optimizers. Across pre-training and fine-tuning, PLUMAGE reduces the optimization gap relative to full-rank training by 36% on average versus the top-k LGE used by GALORE, at essentially the same runtime (within ∼ 1%). Moreover, PLUMAGE reaches GALORE’s terminal loss in fewer steps, including a 30% reduction in 1B-scale pre-training while reusing the full-rank learning rate. Beyond the Adam-based LGE setting, PLUMAGE’s subspace-sampling rule also transfers to the Muon-style SUMO optimizer, reducing its perplexity gap to full-rank Muon by 40% at 130M and 350M scales in our experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.