Fragile by Distillation: Native-Objective Vicinal Alignment for Reliable Synthetic Sets
Abstract
Dataset distillation (DD) plays an important role in creating compact and efficient training sets for reducing downstream computational costs and enabling data sharing. However, previous works have overlooked a critical concern: the condensed datasets are vulnerable to malicious modifications during delivery. Our investigation shows that this weakness is prevalent across representative distillation methods and does not consistently diminish at larger distillation budgets. Models trained on these compromised sets suffer significant performance degradation. To bridge this gap, we introduce Native-Objective Vicinal Alignment (NOVA), a framework which can be incorporated into diverse matching-based DD approaches by the optimization of their native distillation objectives. The key innovation of NOVA lies in jointly matching the synthetic set and its perturbed neighbors to the same real-data target so that the matching holds across a neighborhood of the delivered set rather than at a single point. This matching reduces the shift in the native training signal induced by perturbation. Theoretical analysis and extensive experiments show that NOVA substantially improves the downstream performance of condensed datasets under malicious modifications while preserving clean utility and requiring no changes to the synthetic-set format or the downstream training pipeline. Across seven DD methods, five datasets, and three budgets, NOVA improves accuracy under different perturbations unseen during distillation in 96% of the evaluated settings, with limited impact on clean accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.