GTS: Gap-Targeted Synthesis for Budget-Efficient Synthetic Data Scaling
Abstract
Diffusion models have made synthetic images a practical additional source of training data for visual recognition. In this scaling context, dataset distillation compresses a real training set into a small synthetic replacement, an objective misaligned with settings where the complete real dataset is already available. We instead ask how much accuracy a minimal addition of synthetic images can gain when added on top of that complete real set. We introduce Gap-Targeted Synthesis (GTS), a data-centric framework that diagnoses where synthetic data are needed before generating them. GTS constructs two complementary structural references solely from real features: a mode gap is an observed class mode whose coverage is insufficient; a variation gap is a class-specific interval missing along variation directions shared across classes. GTS converts these deficits into class- and target-level generation quotas and guides a pre-trained diffusion model toward the assigned feature targets during sampling. Gap diagnosis and synthesis require no downstream student classifier, so the same synthetic set can be reused across training architectures. Neither stage draws from an oversized candidate pool. On ImageWoof2, the reported fixed-budget comparisons show higher mean top-1 accuracy than unguided sampling across three student architectures, while paired-noise measurements show that guidance reaches the assigned mode and variation targets more often.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.