OrthVMS: Orthogonal Variational Multiscale Correction for Full-Rank Updates with Low-Rank Optimizer States
Abstract
Training large language models (LLMs) incurs substantial memory overhead, motivating methods that compress optimizer states via low-rank gradient projection. However, projection-based optimizers restrict each update to a compact subspace, discarding the orthogonal gradient component, while maintaining full-size moment states for this residual remains memory-prohibitive. To harness this orthogonal information without sacrificing efficiency, we propose **OrthVMS**, an **Orth**ogonal **V**ariational **M**ulti**S**cale correction framework. Drawing inspiration from variational multiscale modeling (VMS) in computational mechanics, OrthVMS treats the projected space and its orthogonal complement as coarse and fine spaces, respectively. It constructs a full-space preconditioner from two decoupled components: the coarse component maintains compact adaptive moments within the projected space, while the fine component preconditions the orthogonal residual using its historical energy statistics. This residual-driven correction requires only a single persistent vector of additional states. Extensive experiments on LLM benchmarks demonstrate that OrthVMS consistently outperforms existing memory-efficient training methods. Our code is available [here](https://anonymous.4open.science/r/Orthogonal-VMS-C26C).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.