Leaderboard Fidelity: Auditing Synthetic Releases of Private Financial Data
Abstract
Regulated financial data cannot be shared, so synthetic releases are offered as public stand-ins for private fraud-detection benchmarks. A release is used to choose a model, yet it is judged by distributional fidelity and train-on-synthetic accuracy, which do not say whether the models that win on it also win on the private data. We introduce leaderboard fidelity: the rank agreement between the detector leaderboard a release induces and the leaderboard on the private data, calibrated against what resampling the private test period alone produces. Using transfers from a national interbank clearing network, we build a leakage-free benchmark and evaluate sixty releases, five from each of twelve generators, against eleven detectors under equal budgets. Among generators that do not collapse, rank agreement varies more between releases of one generator than between generators, and the releases that reach the calibrated band come from five different generators. None of the distributional criteria we test, nor a train-on-synthetic score, picks the better release of a generator distinguishably better than chance. Measuring rank agreement against the private data does, and only the data holder can do so. The choice carries over to held-out halves and to the later months of the private test period, so a holder can audit several releases and publish the best with the value it scored. When duplicated records cross a random split, most detectors score near perfectly while their order agrees only weakly with the private one. We publish the sixty releases, the protocol, and the code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.