acceptodds
Under review as a conference paper at ICLR 2027

MUQuant: Mask State Adaptive Dual-Center Quantization for Diffusion Language Models

Abstract

Block-diffusion language models repeatedly process partially masked blocks during generation. Within a decoding step, masked and unmasked tokens can have distinct activation centers, making a shared quantization origin poorly suited to either population. Calibration on fully unmasked text also misses the activation mixtures encountered during denoising. **MUQuant** addresses these two sources of W4A4 error with a state-conditioned representation. **MuZero** estimates separate centers for masked and unmasked tokens, and **MuCal** calibrates weights on the reconstructed activations collected along denoising rollouts. **MuEngine** then computes the center and residual contributions in one packed four-bit multiplication, avoiding separate weight passes. Across three diffusion language-model families, this approach improves quantized accuracy while preserving the decoding procedure. On Nemotron-Labs-Diffusion-8B, the evaluated NVFP4 implementation achieves up to **3.24× speedups over BF16**.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.