DyMPQ: Dynamic Mixed-Precision Quantization for Efficient Diffusion Transformers
Abstract
Diffusion transformers (DiTs) have demonstrated strong performance in image and video generation, but their large model sizes and iterative denoising incur substantial computational and memory costs. However, timestep-dependent activation outliers and channel-wise magnitude disparities hinder accurate 4-bit quantization. To address these challenges, we propose Dynamic Channel-wise Mixed-Precision Quantization (DyMPQ) for DiTs. It combines a shared W4A4 main branch with stage-specific high-precision branches to absorb timestep-varying activation outliers. We then apply trajectory-based channel reordering to reduce within-group magnitude disparities in the W4A4 main branch. Building on this mixed-precision representation, we further introduce cross-stage joint calibration to reduce weight quantization error and kernel fusion to enable efficient inference. Extensive experiments on PixArt-, Z-Image, and SANA-Video demonstrate that our method effectively preserves image and video generation quality under low-bit quantization. On an NVIDIA RTX 4090, our method achieves up to single-step speedup and reduces peak GPU memory by up to on Z-Image compared with BF16.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.