acceptodds
Under review as a conference paper at ICLR 2027

BMoEQ: Binary-Coded MoE Quantization with Runtime-Adaptive Mixed-Precision Support

Abstract

Mixture-of-Experts large language models (MoE-LLMs) scale model capacity through sparse activation but retain a substantial memory footprint from expert weights. Existing expert-level mixed-precision methods alleviate this cost by fixing bit-widths according to calibration-based importance estimates, which may not generalize across inputs, domains, and tasks. We propose BMoEQ, a quantization method that combines uniform ultra-low-bit base precision across experts with input-adaptive execution precision through binary-coded quantization (BCQ). To improve quantization accuracy without static expert-level mixed-precision allocation, BMoEQ incorporates saliency-aware error reconstruction, assigning an additional bit plane to a small subset of salient weights while maintaining a common base bit-width across experts. At runtime, online importance scores guide the selective omission of bit planes for less important activated experts, reducing memory traffic and computation without removing their contributions entirely. A single stored representation thus supports multiple execution precisions without requantization or separate lower-precision copies, while preserving access to each expert’s full stored precision. We further develop hardware-efficient GPU kernels that selectively load and process the required bit planes and fuse the salient-weight correction path to minimize latency overhead. Experiments show that the saliency-aware quantization in BMoEQ achieves superior overall accuracy in both 2-bit and 3-bit settings, outperforming expert-level mixed-precision baselines at comparable effective bit-widths. With runtime precision adaptation and optimized GPU kernels, BMoEQ achieves up to 6.4 speedup over the FP16 cuBLAS baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.