Prefix-Point: An Adaptive Non-Uniform Format for Diffusion Model Quantization
Abstract
Post-training quantization of diffusion models is challenging because activation distributions vary substantially across layers and denoising steps. Prior work mitigates this with step-aware scaling for INT quantization, but scaling adjusts only the range of a distribution and not its shape, so it cannot follow activations that shift between heavy-tailed shapes, with most values near zero, and light-tailed shapes spread evenly across the range. Non-uniform formats represent such shapes better, yet existing ones fix a single codebook at design time from an assumed distribution and reuse it throughout the network and therefore degrade generation quality sharply at low precision ( bits). We propose Prefix-Point (PXP), a format that divides the value range into exponent levels and sets the resolution of each level independently, allocating bit-width where it reduces quantization error most, and fitting this allocation separately to the distribution of each layer and denoising step. Since the resulting design space grows exponentially with bit-width, we cast level allocation as a multiple-choice knapsack problem and solve it exactly by dynamic programming. Its floating-point-like structure keeps decoding and the multiplier cheaper than in existing non-uniform formats. Across class-conditional, text-to-image and video diffusion, PXP lowers FID by up to over the strongest existing non-uniform format while being up to faster.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.