acceptodds
Under review as a conference paper at ICLR 2027

MAQuant: Modality-Aware Quantization for Omnimodal Large Language Models

Abstract

Omnimodal large language model (OLLM) deliver strong capabilities but incursubstantial memory and computation costs. Low-bit quantization is a practical remedy, yet existing methods mainly target Large Language Model(LLM) or Large Vision Language Model(LVLM) with modality-agnostic objectives, failing to handle heterogeneous activation distributions, outlier-dominated degradation, and cross-modal error coupling in OLLM. We propose MAQuant, a modality-aware post-training quantization framework that combines Modality-Specific Learnable Transformation (MSL), Exponent-weighted Activation Penalty (EAP), and Cross-modal Error Exchange (CME) to learn modality-specific quantization friendly spaces, protect salient activations, and share channel-wise sensitivity across modalities. EAP and CME are calibration-only objectives and add no inference time loss computation. Experiments on diverse benchmarks and low-bit settings show that MAQuant consistently outperforms strong LLM and LVLM quantization baselines. On Qwen2.5-Omni-7B, MAQuant reduces storage by 68.6% under W4A4 while retaining 97.4% of full-precision average performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.