acceptodds
Under review as a conference paper at ICLR 2027

QSLR: A Quantized + Sparse + Low-Rank Decomposition of LLM Weights

Abstract

Post-training Quantization (PTQ) of Large Language Models (LLMs) is a popular approach to reduce inference costs but can be restrictive under aggressive compression constraints. Hence, recent methods augment quantized weights with (i) sparse corrections to capture outliers (e.g., SpQR is a specialized solver for the quantized-plus-sparse format) or (ii) low-rank corrections to capture shared structure. We combine both in a unified quantized-plus-sparse-plus-low-rank representation, ,and propose an ADMM-based algorithm that fits all three components jointly. Our novel ADMM framework builds upon computationally efficient subroutines and is agnostic to the quantization format. Extensive experiments demonstrate our joint optimization improves upon sequential fitting procedures. reduces the WikiText-2 perplexity of a Llama-3.2-1B from to compared to fitting SpQR followed by a rank-16 correction, at the same compression budget. Additionally, we show that splitting the allocation between sparsity and low rank can improve performance over using either correction alone. For Qwen3-4B with (sym 2-bit) quantization at 2.47 bits per weight, using to fit all three allocations, the mixed representation achieves a WikiText-2 perplexity of 9.76, reducing the perplexity gap to the dense model by 30.9% relative to and 35.2% relative to . Finally, we design GPU kernels that support and establish speedups over FP16 dense computation in matrix-multiplication and decoding benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.