ReAlignQ: Post-Compression Refinement for Joint Sparsification and Quantization of LLMs
Abstract
Quantization and sparsification compress large language models by restricting numerical precision and nonzero support, respectively. Support selected in continuous weight space may be poorly matched to the values that its retained coordinates can realize on the quantization grid. Moreover, criteria applied to individual coordinates overlook interactions among simultaneous changes in support membership and quantized values. We therefore propose ReAlignQ, a refinement framework that jointly selects support exchanges and legal discrete values for incoming weights under fixed quantization scales and a prescribed sparsity budget. For each linear weight matrix, ReAlignQ models interactions among screened candidates using an empirical gradient Gram surrogate built from gradients of individual calibration samples. It retains the resulting gradient signatures only for screened candidates, which reduces the storage needed to model their interactions. We further implement sparse and quantized operators with exact bitmap masks. These operators execute refined nonuniform masks without changing them. We also provide theoretical results that characterize support selection over legal codes, bound the empirical gradient Gram interactions omitted by diagonal screening, and establish surrogate descent while preserving sparsity for accepted exchanges. Across Llama and Qwen models and multiple sparsity levels, ReAlignQ consistently improves language-modeling quality without retraining, changing parent scales, or relaxing sparsity budgets. On Llama 3 8B with an EIWR W8A16 parent at 70% sparsity, it reduces WikiText-2 perplexity from 35.48 to 25.57 and improves mean zero-shot accuracy by 2.6 percentage points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.