I-DiT: Integer-Only Quantization for Diffusion Transformers
Abstract
Diffusion transformers achieve remarkable generative performance but remain challenging to deploy on resource-constrained hardware due to their high computational and memory costs. These devices often rely on specialized AI accelerators whose efficiency comes primarily from low-precision integer-only datapaths. However, existing diffusion quantization methods still rely heavily on floating-point operations, and integer-only frameworks are not readily applicable to modern diffusion transformers due to fixed quantization scales, lack of tiled attention support, and sensitivity to activation outliers. For this, we present I-DiT, the first framework for fully integer-only inference of diffusion transformers without costly fine-tuning. First, we propose a dynamic scale adjustment method for integer-only frameworks, that adapts to varying activation ranges during runtime. Second, we introduce log2-based online softmax that enables numerically stable and memory-efficient tiled attention in integer. Lastly, we reformulate dyadic requantization to support finer-grained activation quantization, localizing the impact of outliers while preserving integer-only execution. Experiments across PixArt, SD3, and FLUX demonstrate that I-DiT preserves generation quality while enabling efficient fully integer-only inference for large-scale diffusion models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.