ModTrain: Training Modulated Diffusion Models for Extreme Low-Bit Quantization
Abstract
Diffusion models have redefined generative AI, but their multi-step inference remains a critical bottleneck for efficient deployment. Quantization offers a powerful remedy, yet extremely low-bit activation quantization remains challenging even for state-of-the-art quantization-aware training (QAT), as its errors are difficult to absorb during training. Modulated diffusion models offer a promising path by exploiting temporal redundancy across denoising steps, but have so far been limited to post-training quantization. We present ModTrain, a training framework for modulated diffusion models. Training these models is non-trivial because historical states couple consecutive denoising steps, complicating data construction, gradient computation, and activation alignment. ModTrain addresses these issues with sequential forward sampling for efficient data construction, history-aware backpropagation for correct gradient flow, and full-precision alignment to mitigate activation distribution shift. ModTrain is orthogonal to existing QAT methods and integrates seamlessly with them. Experiments on LSUN and ImageNet show that ModTrain outperforms post-training MoDiff, reducing activation precision from 4 to 2 bits, and consistently improves QAT baselines at extremely low precision, with up to a 10.90 FID improvement. ModTrain advances extremely low-bit inference and may further motivate hardware support for such computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.