acceptodds
Under review as a conference paper at ICLR 2027

Restoring Minority Generation in Low-Bit Diffusion Models with a Plug-in Repair

Abstract

Efficient deployment of diffusion models has motivated extensive research on low-bit quantization, which can substantially reduce memory and computation while maintaining generation quality under standard sampling. However, we argue that this preservation is not guaranteed to extend to the task of minority generation, which specifically targets generation from low-density but valid regions of the data distribution. We show that quantization error becomes significant in clean estimates at high noise levels, on which existing minority samplers heavily rely yet standard sampling places only a small weight. Based on this analysis, we propose a lightweight plug-in repair that aligns the high-noise clean estimates of quantized models with those of their full-precision counterparts. Our repair combines selective bit allocation with a quantization-consistent low-rank adapter that remains on the deployed quantization grid, and is activated only when computing these estimates, leaving the ordinary sampling path unchanged. Experiments across diverse diffusion architectures, bit widths, and minority samplers show that our repair substantially improves minority generation with negligible overhead while retaining the efficiency of low-bit inference. Code is available here.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.