STRIM: STRucture-aware Low-rank Inference Model
Abstract
Low-rank model compression has become an effective strategy for reducing the size of LLMs while maintaining performance. This paper introduces STRIM (STRucture-aware Low-rank Inference Model), a method for the low-rank compression of large language models. We start from a guiding block-level reconstruction view and derive a greedy, student-input-aware compression procedure. Compared with matrix-wise low-rank compression methods and recent student-input-aware approaches, STRIM explicitly accounts for compression-relevant architectural operators such as residual connections and normalization layers, which helps mitigate cascading compression errors. Experimental results show that STRIM consistently improves task accuracy by over 5% on six classification datasets (e.g., OpenbookQA, WinoGrande, HellaSwag) across a range of model families (e.g., Mistral, Gemma, Qwen) and compression ratios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.