Loss Aware Truncation with Gradients for Low-Rank Decomposition of Large Language Models
Abstract
The continued scaling of large language models (LLMs) has resulted in rapidly growing memory requirements. Singular Value Decomposition (SVD) provides a principled approach to reducing this burden through low-rank compression. However, existing state-of-the-art SVD-based methods have not fully characterized the relationship between singular components and the performance degradation induced by compression. In this work, we show through the first-order Taylor expansion that the local gradient of each layer provides a direct connection between singular value importance and the resulting loss degradation of the compressed model. Building on this insight, we propose Loss-aware Truncation with Gradients (LTG), which transforms the weight space according to local loss sensitivity. This transformation redistributes the singular spectrum such that components with greater influence on the model objective are preferentially retained during truncation. We demonstrate the effectiveness of LTG over related state-of-the-art approaches across three LLMs families and standard benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.