acceptodds
Under review as a conference paper at ICLR 2027

I-DiT: Integer-Only Quantization for Diffusion Transformers

Abstract

Diffusion transformers achieve remarkable generative performance but remain challenging to deploy on resource-constrained hardware due to their high computational and memory costs. These devices often rely on specialized AI accelerators whose efficiency comes primarily from low-precision integer-only datapaths. However, existing diffusion quantization methods still rely heavily on floating-point operations, and integer-only frameworks are not readily applicable to modern diffusion transformers due to fixed quantization scales, lack of tiled attention support, and sensitivity to activation outliers. For this, we present I-DiT, the first framework for fully integer-only inference of diffusion transformers without costly fine-tuning. First, we propose a dynamic scale adjustment method for integer-only frameworks, that adapts to varying activation ranges during runtime. Second, we introduce log2-based online softmax that enables numerically stable and memory-efficient tiled attention in integer. Lastly, we reformulate dyadic requantization to support finer-grained activation quantization, localizing the impact of outliers while preserving integer-only execution. Experiments across PixArt, SD3, and FLUX demonstrate that I-DiT preserves generation quality while enabling efficient fully integer-only inference for large-scale diffusion models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.