acceptodds
Under review as a conference paper at ICLR 2027

In Search of Lost Consistency in Dataset Distillation

Abstract

Dataset Distillation (DD) condenses large datasets into compact synthetic ones, promising comparable performance at a fraction of the training cost. This would make hyperparameter optimization and neural architecture search feasible even for increasingly large foundation models. Recent DD methods report impressive accuracy, yet a critical question remains: Does dataset distillation truly deliver on its promises? We argue that accuracy alone is insufficient; equally crucial is consistency, the extent to which distilled data rank hyperparameters or architectures as the full dataset does. To study this, we introduce the Dataset Distillation Testbed (DataDistillBed), which spans 13 DD methods, five datasets up to ImageNet-1K, and diverse hyperparameter configurations and neural architectures. Our results show that DD delivers on its promises only in part. At matched training budgets, distilled data do not yet replace the full dataset. As proxies for search, however, they rank training hyperparameters largely as the full dataset does (median Spearman correlation 0.82 at 50 images per class), though architectures less faithfully (0.52). Case studies that use DD to retrace the design of ConvNeXt and EfficientNet show its practicality and limitations. We further analyze how consistency relates to accuracy, compression ratio, training cost, and individual hyperparameters. Building on these findings, we design a search strategy that uses distilled data as a cheap first stage and matches the strongest full-data baselines at a third to half the cost. Because accuracy depends strongly on training hyperparameters, we also propose an evaluation protocol that compares DD methods under shared, randomly drawn training configurations and reports consistency alongside accuracy. We hope that the release of DataDistillBed will catalyze more rigorous, reproducible, and practically effective DD methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.