Spectral Gradient Surgery for Domain Generalizable Dataset Distillation
Abstract
Dataset Distillation (DD) synthesizes a compact synthetic dataset that preserves the training utility of a full dataset. However, its standard formulation assumes that test data follow the same distribution as training data, which rarely holds in practice. A straightforward extension—applying post-hoc Domain Generalization (DG) techniques to distilled data—is ill-suited because existing DG methods rely on the natural diversity of real datasets, which compact synthetic sets inherently lack, and their augmentation overhead conflicts with the efficiency goal of DD. We therefore study Domain Generalizable Dataset Distillation (DGDD), a setting in which distilled datasets are evaluated on target domains unseen during distillation, through the lens of Distribution Matching (DM). We hypothesize that the OOD vulnerability of DM stems from compressing domain-consistent and domain-specific information indiscriminately, and propose Spectral Gradient Surgery (SGS). SGS computes domain-wise DM gradients with respect to the synthetic images, transforms them into the Fourier domain, and measures per-frequency phase agreement across source domains. It then augments the standard DM update with two terms: a class-relevant component that emphasizes frequency components consistent across source domains, and a domain-relevant component that preserves source-domain variation to diversify the synthetic set. Across three benchmarks of increasing resolution, SGS consistently improves the average OOD accuracy of two DM methods without modifying their objectives or downstream training cost, at the price of additional offline distillation cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.