Factorized Chain Utility: A Calibrated Divergence Diagnostic for Synthetic Tabular Data
Abstract
Synthetic tabular data helps share information while protecting commercial or personal interests, so its value depends on how well it replicates the joint distribution of the real data. Machine-learning utility measures whether a model trained on synthetic data predicts a target variable on real data. Because this score depends only on a single conditional distribution, a synthetic table can score well on this measure while failing to preserve important relationships among the other variables. We propose Factorized Chain Utility (FCU), a diagnostic for synthetic tables based on a sparse chain of conditional prediction tasks. FCU uses proper log scores to quantify distortion in selected conditional relationships. The selected conditionals multiply to a distribution, unlike the leave-one-out conditionals used by prior all-column metrics. This gives FCU's aggregate score a well-defined divergence interpretation for the selected chain. We report a pruning deficit that quantifies the dependence information omitted by the sparse factorization, and per-factor scores identify the relationships responsible for degraded utility. Controlled validation shows that FCU responds to dependency loss even when marginal distributions are preserved. Across diverse datasets and generators, FCU further separates synthetic datasets that appear comparable under single-task utility or discriminator-based evaluations, providing a principled generator ranking with actionable diagnostics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.