acceptodds
Under review as a conference paper at ICLR 2027

SparseSliderQuant: Sparse-Quantized LLM with Cross-Block Error Coordination

Abstract

Sparse quantization is an attractive compression target for Transformer decoders because it simultaneously reduces numerical precision and arithmetic density. However, aggressive configurations such as W4A4 quantization with 2:4 structured sparsity expose a fundamental gap between local compression and cross-block consistency. To bridge this gap, we propose SparseSliderQuant, a unified rewrite framework that coordinates sparsification and quantization errors across Transformer blocks. SparseSliderQuant first employs sliding group alternation to iteratively optimize sparse masks and quantization scales within overlapping block windows, enabling the two compression processes to adapt to each other rather than being optimized independently. It then performs dual-source mismatch calibration by explicitly decomposing pruning and quantization residuals, accumulating their mismatch in a residual buffer, and redistributing it to the remaining unpruned weights. Finally, cross-block gradient anchoring uses next-block output drift to jointly guide mask and scale updates, aligning local reconstruction with downstream behavior throughout the decoder. SparseSliderQuant introduces no auxiliary modules or additional inference-time overhead. Experiments on LLaMA and Qwen models show that it consistently preserves language-model performance across model scales and a broad range of language understanding and reasoning tasks under W4A4 quantization with 2:4 structured sparsity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.