SLORR: Simple and Efficient In-Training Low-Rank Regularization
Abstract
Direct spectral regularization is a natural way to induce low-rank weights, but common objectives are costly at modern scale because they typically require repeated singular value decompositions (SVDs). Existing alternatives reduce this cost through reformulations that introduce architectural reparameterization, cached spectral state, predefined target ranks, or periodic decompositions, thus hindering practicality. We introduce SLORR to make direct spectral regularization practical without these restrictions while retaining competitive post-compression performance. Our SVD-free, stateless, architecture-preserving framework uses GPU-friendly approximations of the spectral quantities needed for regularizer forward and backward passes. SLORR supports different spectral objectives, instantiated here with Hoyer spectral sparsity and nuclear norm regularization. We evaluate induced compressibility and performance after post-training low-rank compression across different variants, using ImageNet-1K pretraining and continued training of ResNets and Vision Transformers, and LLM pretraining at 135M, 560M, and 1B parameters. Across vision settings and at 135M LLM scale, SLORR achieves competitive compression–quality tradeoffs with recent strong low-rank regularizers, often with greater robustness across settings, while retaining low training overhead and avoiding their limitations. We further scale SLORR to 560M and 1B LLMs: it substantially improves post-compression performance over matched unregularized training. These gains persist across different language model token budgets and four optimizers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.