acceptodds
Under review as a conference paper at ICLR 2027

BitMoE: Post-Training Binarization for MoE-based Large Language Models

Abstract

Large language models (LLMs) with Mixture-of-Experts (MoE) architectures have achieved remarkable progress, yet their substantial weight storage and memory traffic hinder practical deployment. Weight binarization offers extreme compression, but transferring dense-model methods to MoE faces three coupled challenges: decomposition–binarization mismatch, interaction-dependent residual allocation, and quantization-induced expert-shift. To address these challenges, we propose BitMoE, a post-training binarization framework coordinating shared representation, residual capacity, and routing stability. First, Routing-Coupled Binary Factorization alternately refines higher-precision shared factors and expert-specific binary factors to reduce the reconstruction error of routed expert aggregation. Second, Conditional Expert Residual Allocation selects additional binary residual blocks by their conditional reconstruction gain per added storage bit after readapting the shared output factor, updating gains as corrections accumulate. Third, Margin-Aware Optimal Brain Routing uses boundary-aware curvature to preserve selected–unselected expert margins and mitigate expert-shift. BitMoE retains all experts and performs optimization offline under an explicit storage budget. Extensive experiments demonstrate that BitMoE consistently outperforms state-of-the-art binary baselines across multiple MoE models and benchmarks. For instance, on Qwen3-30B-A3B, BitMoE delivers 54.2% lower perplexity and improves average zero-shot accuracy by 19.71 percentage points over the strongest binary baseline, alongside an over 2 inference speedup over the BF16 model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.