Simultaneous Finite-Sample Certification of Synthetic Data Generators under Fidelity Constraints
Abstract
Synthetic data generators (SDGs) are increasingly used in high-stakes settings, yet practitioners often lack a principled way to determine whether the generated data adequately represent the real data distribution for a specific dataset. While theoretical guarantees for SDGs do exist, they are often asymptotic and therefore do not directly provide finite-sample guarantees for the trained generator and dataset at hand. As a result, it can be difficult to determine which among several candidate generators satisfy a required standard of data fidelity. We address this problem by formulating the certification of SDGs as a multiple hypothesis testing problem over a collection of candidate generators. For each generator, we test whether a population-level fidelity criterion, defined through a user-specified discrepancy measure between the real and synthetic distributions, falls below a prescribed tolerance level. By constructing valid tests from finite-sample concentration bounds and applying a family-wise error rate controlling procedure, our framework guarantees, with probability at least , that all selected generators satisfy the specified fidelity criterion. This provides dataset-specific, finite-sample certification of synthetic data fidelity and allows practitioners to screen multiple candidate generators and retain a statistically supported subset that meets a desired fidelity standard before considering other downstream objectives. The framework is agnostic to the generator architecture and applies to a broad class of discrepancy measures whose estimators admit suitable concentration inequalities, making it useful for applications where reliable finite-sample assessment of synthetic data is important.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.