acceptodds
Under review as a conference paper at ICLR 2027

ASQ: An Improved Approach to Unifying Sparsification and Quantization of LLMs

Abstract

Sparsification and quantization are widely adopted for the efficient deployment of large language models (LLMs). However, they are often optimized independently, overlooking their intricate interactions and yielding suboptimal compression trade-offs. To address this, we propose ASQ, an ADMM-based framework that jointly optimizes sparsification and quantization under a unified formulation, supported by theoretical convergence guarantees. At its core, ASQ admits an exact, closed-form joint projection that concurrently resolves sparse support and quantized values by accounting for both their importance to the objective criterion and their representability on a quantization grid. Across multiple model architectures and scales, ASQ substantially outperforms existing joint compression methods under aggressive compression, with the performance gap widening significantly under extreme sparsity and low-bit regimes. Moreover, under hardware-compatible 2:4 sparsity with INT4 quantization, ASQ achieves a 4.78 memory reduction and up to 3.20 end-to-end decode speedup on Qwen3-8B-Base relative to dense BF16 inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.