SPARCQuant: Sparse Adaptive Repair of Critical Channels for MLLMs Quantization
Abstract
Post-training quantization (PTQ) for multimodal large language models is challenging because heterogeneous modalities induce distinct activation distributions in the shared decoder. Smoothing-based quantization has proven effective and is widely adopted for large language model quantization, yet directly extending it to these models can cause smoothing factors estimated from mixed calibration data to be dominated by high-magnitude modalities, thereby over-smoothing other modalities and increasing activation quantization error. Against this backdrop, the analysis in this paper reveals a previously overlooked structural property: **high sensitivity to modality-specific smoothing mismatch is concentrated in a small subset of channels**. Motivated by this finding, this paper proposes **SPARCQuant**, a sensitivity-guided framework for **sparse adaptive repair of critical channels** in multimodal large language model quantization. SPARCQuant selects a layer- and module-adaptive base modality to define shared smoothing and ranks channels using robust activation-range discrepancy and activation-magnitude importance. A coarse-to-fine proxy reconstruction search with online scheduling then **automatically selects the sparse channel support and its size**, adapting the repair to each layer, module, and non-base modality **without manual channel selection or per-layer budget tuning**. Selected-channel ridge compensation corrects residual errors while preserving a single shared quantized weight. Experiments on vision-language and omni-modal models demonstrate **improved robustness under low-bit quantization**. A CUDA implementation with custom kernels enables efficient deployment, achieving **2.66×–2.87× end-to-end inference speedup over BF16**.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.