Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts
Abstract
Mixture-of-Experts (MoE) architectures allow frontier language models to scale toward trillions of parameters, but their deployment is constrained by massive memory footprints and bandwidth limitations. While modern accelerators feature Sparse Tensor Cores (SpTCs) to reduce weight storage and boost throughput via low-precision semi-structured sparsity, exploiting them in MoEs is hindered by severe model-quality degradation and the lack of grouped sparse GEMM primitives. We present an end-to-end hardware–software co-design framework that compresses expert weights into hardware-native, low-precision sparse representations and accelerates their execution on SpTCs. Algorithmically, our framework relaxes discrete semi-structured support selection through continuous reparameterization, enabling differentiable joint optimization with quantized weights under a router-weighted reconstruction objective scaled through expert-parallel compression. Systemically, we implement a custom grouped sparse GEMM kernel tailored to sparse, low-precision MoE inference on SpTCs. Across MoE scales ranging from 30 billion to frontier-scale trillion-parameter models, our framework advances state-of-the-art joint sparse quantization by up to percentage points in task accuracy while preserving of the original models' performance. System-level benchmarks on NVIDIA B200 GPUs show that our kernel outpaces the vendor baseline by up to , delivering higher serving throughput and up to a reduction in end-to-end latency. These results establish hardware–software co-design as a practical and necessary pathway toward scalable, highly efficient MoE deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.