acceptodds
Under review as a conference paper at ICLR 2027

Rotate All, Store Few: Decoupling Coverage from Storage for Memory-Efficient Training

Abstract

Adam improves when the gradient is normalized in its singular basis, which we call . Full-rank rotation exceeds the memory of Adam by storing a basis and states for every direction, whereas low-rank gradient projection rotates only the top- directions but degrades sharply as decreases. This is because a single rank sets both the , the number of rotated directions, and the , the number of directions with optimizer states. The two decisions nonetheless need not be coupled. The rotation coverage should be full since the whitened update magnitude stays nearly flat along the spectral index, whereas the state storage can be reduced since adjacent second moments have high cosine similarity. We, therefore, propose , which decouples rotation verage nd state orage. rotates every direction but maintains optimizer states for only stored directions and estimates the updates of the rest from their stored neighbors. On LLaMA pre-training, outperforms Adam with – Adam's optimizer memory. It also improves on low-rank methods at similar memory. Halving the storage budget changes perplexity by less than 1% at every scale.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.