FROM PREDICTABLE DISCREPANCY TO RESIDUAL VARIANCE: CORRECTING QUANTIZED DDIM SAMPLING
Abstract
Post-training quantization (PTQ) reduces diffusion-model inference cost, but denoiser-output discrepancies perturb subsequent sampling states. We introduce State-Conditioned Predictive Error Correction (SPEC), a two-stage output corrector for a fixed quantized denoiser under DDIM-family sampling. At sparse refresh steps, its temporal branch learns an output-matching offset for the quantized denoiser's time condition from full-precision supervision on the same latent state; the DDIM grid and transition coefficients remain nominal. A full-resolution branch is then trained on trajectories recollected with this adjustment active to predict the remaining same-state output discrepancy. On held-out trajectories, the Predictable Error Ratio (PER) shows that SPEC reduces the summed squared same-state output discrepancy between the quantized and full-precision denoisers by 82.7%–88.1% across the evaluated bitwidths. For stochastic DDIM, Student-\(t\)-estimated variance reallocation (\(t\)-VR) models the remaining post-SPEC residual variance and reallocates the nominal stochastic variance between its estimated contribution and the explicit Gaussian innovation. With 100-step stochastic DDIM on CIFAR-10 at W4A8, FID falls from 6.84 for Q-Diffusion to 5.83 for SPEC and 5.66 with \(t\)-VR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.