acceptodds
Under review as a conference paper at ICLR 2027

Not All Tokens Are Equal: Rethinking Quantization for Diffusion Language Models

Abstract

Quantization is a standard approach for large language models' deployment. Most existing quantization methods for large language models are designed for autoregressive models and the quantization objective is based on layer-wise mean square reconstruction error. In this paper, we show that such objective is mismatched to diffusion language models. Unlike autoregressive models, diffusion language models distinguish between visible tokens, which provide conditioning context, and masked tokens, which are the actual prediction targets. This structural asymmetry induces a fundamentally different quantization objective. Our analysis shows that minimizing the standard layer-wise mean square reconstruction error can yield a loose bound for quantization error in diffusion language models. Motivated by the above analysis, we explicitly disentangle visible and masked tokens as two groups and propose MV-Quant, which minimizes the worst-group reconstruction error. We further prove that, for an n-layer Transformer, MV-Quant yields a tighter quantization error bound than the standard MSE objective. Extensive experiments across multiple diffusion language models and multiple benchmarks validate the effectiveness of our proposed MV-Quant.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.