acceptodds
Under review as a conference paper at ICLR 2027

DeDD: Decontaminating Dataset Distillation under Label Noise

Abstract

Dataset distillation (DD) compresses a large training set into a small synthetic set on which models can be trained to achieve accuracy comparable to that obtained on the original data. Most existing DD methods, however, assume that accurate gold labels are readily available, whereas annotations of large-scale datasets are inevitably corrupted by noise. Under such label noise, existing DD methods degrade substantially, even when preceded by an additional label-cleaning stage. To address this problem, we propose Decontaminating Dataset Distillation (DeDD), a framework that distills a clean synthetic set directly from label-corrupted training data. By modeling the observed class-conditional distributions as mixtures of their latent clean counterparts and analytically inverting this mixture, our method recovers unbiased estimates of the clean class prototypes and matches the synthetic set to them in a single stage, without explicit label cleaning or sample selection. Theoretically, our objective yields, in expectation, the same optimization gradients as distillation on clean data. Experiments with simulated and real-world label noise demonstrate that DeDD outperforms state-of-the-art DD methods. The code package is available at https://anonymous.4open.science/r/DeDD_ICLR27/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.