acceptodds
Under review as a conference paper at ICLR 2027

Direction-Aware Low-Rank Weight Factorization for Post-Training Compression

Abstract

Low-rank factorization is a widely used post-training compression technique that typically minimizes weight reconstruction error. However, preserving reconstruction quality alone does not necessarily preserve model performance: it describes how much the weights are changed, but not whether the induced perturbation moves the model toward lower-loss or higher-loss regions. In neural networks, this distinction is crucial. Under a local first-order view, the loss variation is determined by the interaction between the loss gradient and the factorization perturbation, indicating that perturbation direction is a key factor overlooked by reconstruction-only methods. Motivated by this observation, we propose a direction-aware low-rank factorization framework that encourages low-rank perturbations to follow descent-aligned directions while maintaining compression constraints. We further develop two practical variants that emphasize either performance preservation or stronger compression. As a post-training factorization method, our approach does not require multi-epoch post-factorization finetuning. Experiments on vision models, NLP models, and large language models show that our method consistently achieves a better performance-compression trade-off than existing low-rank baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.