Slide&Prune: A Unified Pruning Framework for Joint Low Rank Compression and Channel Sparsification via Sliding Window
Abstract
Low-rank compression reduces the deployment cost of large language models (LLMs), but preserving model quality under aggressive compression remains challenging. Conventional top-k singular-value truncation is optimal for the corresponding layer-wise reconstruction objective, yet this guarantee does not extend to the behavior of multiple composed Transformer blocks. We introduce Slide&Prune (S&P), a post-training framework that extends sliding-window reconstruction to low-rank compression, enabling out-of-order singular-component selection according to its effects across neighboring blocks. This formulation generalizes rank truncation into a pruning problem, allowing learnable masks to jointly select singular components, attention heads, and feed-forward network channels under a shared parameter budget. S&P uses a two-stage pipeline of global allocation and sliding-window refinement while keeping pretrained weights and SVD factors fixed. Experiments across OPT, Vicuna, and LLaMA models from 6.7B to 30B parameters establish state-of-the-art compression–quality trade-offs among the evaluated structured pruning and low-rank compression methods, with particularly strong gains under aggressive compression. On LLaMA-1-7B at a compression ratio of 0.4, S&P achieves WikiText-2 perplexity of 13.75, compared with 45.17 for the strongest evaluated baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.