TeleQuant: Temporal Noise Shaping for Post-Training Quantized Diffusion Transformers
Abstract
Post-training quantization lowers the cost of the repeated denoiser evaluations that diffusion Transformers use for image and video generation. Because those evaluations are chained, the rounding errors they leave behind interact across steps. Correlated rounding errors can reinforce, accumulate into structured trajectory drift, and persist in the returned sample. This drift can degrade image and video fidelity. We propose TeleQuant, a training-free temporal noise-shaping framework that feeds a discarded activation-rounding residual into the next evaluation, with a gain matched to adjacent solver responses and activation scales. We derive a first-order telescoping identity: under ideal state storage, accumulated error separates into a terminal residual and a sensitivity-variation term, motivating selective layer placement and damped feedback. INT8 residual state and a fused activation operator add the compensation on the host's path. The feedback lowers accumulated latent drift and improves paired fidelity to matched FP16 outputs. Evaluations on DiT-XL/2, PixArt-, SDXL, and Open-Sora 1.2 show improved image fidelity and video consistency over frozen ViDiT-Q and SVDQuant hosts, with matched rounding ablations supporting the complete gain. The profiled image pipelines retain most of the host speedup, at – latency. Code is available at https://anonymous.4open.science/r/TeleQuant-8217/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.