Second-Order Modality-Expert Joint Quantization for MoE VLMs
Abstract
Post-Training Quantization (PTQ) is critical for compressing memory-intensive Mixture-of-Experts Vision-Language Models (MoE VLMs), yet existing methods suffer from fundamental limitations: they rely on first-order approximations that fail to capture the complex interactions between quantization errors, expert routing, and cross-modal fusion, and they treat vision and language modalities as homogeneous, ignoring their inherent differences in feature sparsity and expert activation frequency. To overcome these limitations, we propose Second-Order Modality-Expert Joint Quantization (SOME-JQ), a PTQ framework that leverages second-order task-loss approximation and dual-aware optimization to minimize quantization error for MoE VLMs. Specifically, SOME-JQ makes key contributions. Second-Order Error Analysis (SOEA) decomposes the total quantization error into router logit distortion, expert output reconstruction error, and cross-modal fusion error, formally establishing that the dominant term is a routing-weighted combination of expert and fusion errors. Modality-Expert Joint Hessian (MEJH) construction integrates token-expert affinity, per-token modality information, and cross-modal fusion weights into an enhanced Hessian matrix that guides calibration to prioritize critical experts and cross-modal interactions. Learnable Modality-Expert Clipping (LMEC) introduces task-aware learnable clipping parameters for each expert–modality pair, jointly optimizing co-activated experts and cross-modal tokens to minimize reconstruction error. Extensive experiments on Kimi-VL and Qwen3-VL across diverse multimodal benchmarks demonstrate consistent improvements over strong PTQ baselines, especially under extreme low-bit settings, where SOME-JQ achieves substantial accuracy gains over the previous state of the art.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.