Beyond Augmentation Performance: A Utility, Attribution, and Privacy Audit for Synthetic Clinical Spectra
Abstract
Generative models for small clinical datasets are commonly evaluated by adding synthetic samples to real training data and measuring downstream performance. This mixed-training benchmark can confound generator quality with class rebalancing, does not isolate the contribution of generated samples, and provides no evidence of patient-level privacy. We address these limitations with a structured empirical audit that evaluates three distinct properties: utility through train-on-synthetic, test-on-real substitution under prespecified equivalence testing; attribution through reverse transfer and zero-signal controls; and privacy through membership-inference and copy-rate gates. On ATR-FTIR, a latent-diffusion generator with near-chance synthetic-to-real transfer matches its corrected counterpart in augmentation AUC, showing that mixed-training performance can conceal generator failure. Across ATR-FTIR (\(n=584\)) and LIBS (\(n=112\)), linear PCA-Gaussian matches or exceeds the individual nonlinear generators, while mixtures improve substitution stability. Post-hoc corrections create apparent utility even from label-shuffled or noise inputs, whereas SMOTE shows strong utility yet suffers from near-perfect membership discrimination under the evaluated distance-based attack. In the end we produce fully synthetic artifacts with a classifier passing equivalence criteria with a one trained on real data. However, a post-hoc full-spectrum attack detects membership signal in the assembled LIBS artifact (\(AUC=0.58\), \(p=0.004\)), showing that empirical privacy depends on the adversarial representation. Thus, augmentation performance alone does not establish generator utility; utility, attribution, and privacy require separate evaluation. Our audit is a falsification framework, not a formal privacy guarantee, clinical validation, or certification of release safety.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.