NMP-SVD: Rethinking Dense Projections via N:M-Structured Low-Rank Compression for LLMs
Abstract
Existing large language model (LLM) compression methods based on Singular Value Decomposition (SVD) mainly focus on weighted reconstruction, singular-value truncation, and rank allocation, while largely leaving the representation of retained projection directions unexplored. Most existing methods preserve these projections as dense factors. We revisit this assumption through controlled perturbation experiments and find that low-rank representations can tolerate substantial directional approximation, with expanded approximate projections often outperforming smaller exact dense representations. Motivated by this observation, we propose NMP-SVD ( Projection SVD), a plug-and-play structured-sparsity enhancement for existing SVD compressors. NMP-SVD directly optimizes both low-rank factors under the original weighted reconstruction objective using dual-factor -OBS calibration, expands projection directions from the residual of the actual sparse solution, and applies block-wise backfitting to correct sequential expansion errors. A per-linear-layer selector further retains dense realizations for sparsity-sensitive layers. NMP-SVD preserves the weighting and rank-allocation strategies of the underlying compressor and is therefore compatible with diverse SVD frameworks. Experiments across multiple models and datasets demonstrate improved quality–compression trade-offs, higher prefill efficiency, and reduced memory usage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.