acceptodds
Under review as a conference paper at ICLR 2027

ULMoE: Algorithm-System Co-design for Efficient Ultra-Low-Bit MoE Compression and Serving

Abstract

Mixture-of-Experts (MoE) models now dominate frontier LLMs. However, the massive parameter count of experts makes them expensive to fit into GPU memory, and the traffic of reading expert weights on every decode step makes them expensive to serve. We show that ultra-low-bit weight quantization addresses both challenges, and that it pays off more in MoE than in dense models, because sparse routing makes the memory-bound region of grouped GEMM wider than dense GEMM, covering the practical serving batch sizes. Achieving practical serving speedup requires the quantization format and the kernel to be designed together. Methods such as vector quantization and expert-level mixed precision are designed for accuracy alone, and although they reach two bits, no grouped GEMM kernel can execute them efficiently. Methods designed around kernels achieve efficient serving performance but stop at four bits. We present ULMoE, an algorithm-system co-design for accurate and efficient ultra-low-bit MoE serving. Its format spends 2.5 bits per weight and places the quantization group scales on two integer levels. The metadata for 64 weights fits in one aligned 32-bit word for efficient memory access. On top of the format we build a fast W2A16 grouped GEMM kernel and integrate it into SGLang. Across seven MoE models from three families, ULMoE compresses expert weights at to effective bits and loses average accuracy points. At matched interactivity it delivers the serving throughput of BF16 on Qwen3.5-35B and that of Int4 on Qwen3.5-122B.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.