HEPH: 3-Bit Grid-Aligned Trellis Quantization of MXFP4 Mixture-of-Experts for Accurate, Hadamard-Free Serving
Abstract
A growing class of mixture-of-experts models, including gpt-oss, Kimi K3, and DeepSeek-V4, ships routed experts natively in the MXFP4 hardware microscaling format after quantization-aware post-training. HEPH is a trellis-coded codec for these weights. Existing methods such as QTIP and QuIP# achieve strong rate–distortion performance but decode to arbitrary FP16 values in a randomized-Hadamard-rotated basis. Serving them requires a Hadamard transform on every forward pass and cannot use the native FP4 tensor-core kernel. HEPH instead targets the hardware grid: it reuses each checkpoint block’s UE8M0 exponent as its exact scale and restricts the trellis codebook to the 15 distinct E2M1 values, making every decoded weight a valid MXFP4 element. This design improves both accuracy and serving speed. On gpt-oss-20b, HEPH achieves **80.03** mean accuracy across five zero-shot benchmarks at 3.04 bits per weight, versus **77.72** for a rate-matched QTIP baseline and within 0.76 points of the 4.25-bpw MXFP4 checkpoint. When both compressed streams remain resident and decode weights in-kernel on each forward pass, HEPH reaches **96.1** tokens/s for batch-1 decode on a Grace Blackwell GPU in SGLang, versus **74.2** for QTIP (); it remains – faster across batches of 1–32. HEPH also supports a second serving mode: its approximately 3-bpw storage representation can be expanded into a resident 4.25-bpw MXFP4 tensor and executed by the unmodified production fused grouped-MoE kernel. This trades resident memory for throughput while retaining the storage saving. In this mode, HEPH reaches **382** tokens/s at batch 1, versus **72.8** for QTIP (), with speedups of – across batches of 1–32. Aligning the codec with the native FP4 grid thus yields both higher accuracy at a matched compression rate and faster mixture-of-experts serving.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.