acceptodds
Under review as a conference paper at ICLR 2027

Configuration-Valued Expert Importance for Mixed-Precision Quantization of Multimodal Mixture-of-Experts

Abstract

Sparse activation does not eliminate the expert-weight storage cost of multimodal Mixture-of-Experts (MoE) models, motivating mixed-precision quantization under memory constraints. However, existing importance-based approaches often rely on isolated expert or projection scores, overlooking the context dependence of expert importance, the natural hierarchy between experts and their internal projections, and the coupled quantization effects within each expert. We propose Config-Owen, a hierarchical game-theoretic framework for mixed-precision MoE quantization. Config-Owen formulates expert-level precision allocation as a cooperative game, uses Owen-style attribution to preserve the expert–projection hierarchy, and incorporates within-expert Harsanyi interactions to capture coupled quantization effects, yielding configuration-valued utilities for precision allocation. We efficiently estimate these utilities and optimize the resulting precision assignment under a global memory budget. Across four multimodal MoE models and Mixtral-87B, Config-Owen consistently improves low-bit performance preservation; on DeepSeek-VL2-S at 2.58 bits, it retains 98.39% of the FP16 five-task average score. Code will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.