BTT: Quantizing Diffusion Language Models via Block-wise Token Transformations
Abstract
Low-bit activation quantization remains challenging for diffusion language models (dLLMs), whose activations vary across token positions and denoising states. Existing channel transformations reshape feature coordinates within tokens, leaving the token dimension available for complementary transformations. We propose Block-wise Token Transformation (BTT) for linear-input activation quantization in dLLMs. BTT applies randomized Hadamard transforms within blocks of simultaneously processed tokens before activation quantization and restores token coordinates after the linear projection. This construction preserves the linear function in exact arithmetic without quantization, while redistributing activations presented to the low-bit quantizer. The block-wise design enables efficient computation and composes naturally with fixed channel rotations. Because token transformations do not benefit all operators equally, we select them independently for each linear operator by minimizing output reconstruction loss on shared calibration inputs with fixed quantized weights, retaining the identity when no candidate improves reconstruction. Experiments show that adding BTT to QuaRot and DuQuant improves downstream accuracy under W4A4 quantization. These results demonstrate the potential of token-dimension transformations as a complement to channel transformations for low-bit dLLM quantization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.