acceptodds
Under review as a conference paper at ICLR 2027

DRAM: Memory-Efficient Muon via Rank Reallocation in Principal-Random Subspaces

Abstract

Memory-efficient Muon methods aim to reduce storage overhead while preserving strong performance, yet existing approaches compromise either convergence guarantees or practical efficiency. In this paper, we focus on subspace optimization for Muon, where the choice of subspace determines the gradient information retained for updating. To preserve dominant information while providing the coverage needed for stochastic convergence guarantees, we construct a principal-random subspace that combines leading singular directions of the stochastic gradient with random directions sampled from their orthogonal complement. Under a fixed rank budget, there is a trade-off between principal exploitation and complementary coverage, making rank allocation important for practical optimization. Our key observation is that the preferred balance shifts from principal-heavy early in training to random-heavy later. Motivated by this shift, we propose DRAM, a memory-efficient Muon optimizer that reallocates principal and random ranks across stages while maintaining momentum and performing matrix-sign normalization entirely in the subspace. Theoretically, we establish convergence guarantees for general rank-allocation schedules with nondecreasing coverage factors, obtaining rates of in the stochastic setting and in the deterministic setting, with fixed allocation as a special case. Using a simple linear reallocation schedule, DRAM achieves the best low-memory performance across three pre-training scales and nine fine-tuning settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.