How Many, Not Which: Automatic Data Budgets for Conditional GANs Using Gradient Spectra
Abstract
Data subset selection promises cheaper training. While the coreset literature focuses on deciding which samples to keep at a practitioner specified size, our evidence suggests that, for conditional GANs, the predominant factor is how many samples. At equal budgets, uniform random selections during training runs tie or outperform gradient matching, coverage matching, and difficulty scoring. Meanwhile, changing data quantity and allocation for each condition moves quality well outside between-seed variation. To address this gap in determining a subset’s quantity, we introduce FLARE (Fast Learning via Adaptive Rank Estimation), which derives the data budget from the effective rank of each condition's centered per-sample critic gradient spectrum, i.e., the number of independent directions the gradient signal spans. That budget is then split across conditions proportionate to those ranks, tilted toward the conditions the generator currently worst fits as ordered by per-condition sliced Wasserstein distance. Because the budget is an output of that spectrum rather than a required input, there is no size parameter to tune. We apply FLARE to optiGAN—a conditional Wasserstein GAN with Gradient Penalty (WGAN-GP) that stands in for state-of-the-art Monte Carlo simulation of optical photon transport in radiation detectors—to determine a coreset size under 2% of a 513k-sample pool. With random subsets at this coreset size, optiGAN trains 13x faster (including incurred selection costs, excluding evaluation) for the same number of optimizer steps, matching full-data training on global sliced-Wasserstein similarity at a minimal quantified cost in per-condition fidelity. Because the criterion is solely based on the critic's own gradients, it transfers across domains without re-tuning. As such, we apply FLARE unchanged to calorimeter shower simulation, and to tabular synthesis on a network-intrusion benchmark. Because FLARE’s selected fraction differs substantially across these domains, no single ratio fits them all; deriving the budget instead of tuning it removes the sweep that data reduction normally demands.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.