acceptodds
Under review as a conference paper at ICLR 2027

MDMixQ: Modality-Decoupled Mixed-Precision Quantization for Multimodal MoE Compression

Abstract

Post-training quantization is widely adopted for reducing the memory footprint of large-scale Mixture-of-Experts (MoE) models, and expert-level mixed precision further improves the accuracy-compression trade-off by allocating heterogeneous bit-widths across experts. In multimodal MoE settings, however, we find that existing mixed-precision pipelines overlook inter- and intra-modality routing non-uniformities that skew expert bit allocation and degrade accuracy under compression. To address these failure modes, we propose MDMixQ, a mixed-precision quantization framework comprising three targeted components: Modality-Decoupled Calibration (MDC), which isolates a clean text-routing signal from multimodal inputs; Salience-Guided Filtering and Rectification, which retains only confident text-side routing events and rectifies sign-inconsistent logits; and Dynamic Mixed-Precision Allocation, which ranks experts by the refined salience scores under a global bit budget and applies modality-aware activation quantization during MoE inference. At a 2.52-bit average expert budget, the best MDMixQ configurations retain 96.29%, 96.05%, and 95.33% of FP16 performance on Qwen3-VL-30B-A3B, Kimi-VL-16B-A3B, and Qwen3-Omni-30B-A3B, respectively, outperforming prior expert-priority methods across two model families. The framework also transfers to the MXFP format with 99.72% performance retained on Qwen3-VL-30B-A3B, confirming compatibility with emerging hardware-oriented quantization standards.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.