Towards Efficient Diffusion Training with Noisy Labels
Abstract
Label noise can disrupt the correspondence between images and class labels, degrading conditional alignment and generation quality in diffusion models. Transition-aware weighted Denoising Score Matching (TDSM) addresses this problem by estimating instance- and time-dependent clean-label posterior probabilities using a label transition matrix and a classifier, and using these probabilities to aggregate class-conditional denoising predictions. However, this mechanism requires a separate denoiser evaluation for each class involved in the aggregation, incurring substantial training overhead when the number of classes is large. To reduce this cost, we propose mixed-label conditioning, which combines learnable class embeddings according to the estimated clean-label posterior probabilities and conditions the denoiser on the resulting mixed representation. By encoding label uncertainty directly in the conditioning input, our method requires only a single forward pass for each mixed-label prediction, avoiding explicit class-wise prediction and output aggregation. To further improve training efficiency and effectiveness, we introduce mixed diffusion-model-guided sample selection, which compares the diffusion losses of mixed-label and reference predictions to select training examples that benefit from conditioning at the sampled noise level. Experiments on CIFAR-10 and CIFAR-100 under noisy-label settings evaluate generation quality, conditional alignment, and computational cost. The results demonstrate generation performance comparable to that of TDSM with substantially reduced training overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.