Tabular Pancakes: Cryptographic Watermarks Against Dataset Fabrication
Abstract
Tabular generative models can fabricate and impute scientific datasets, bypassing fraud forensics and heightening malpractice risks. Since no statistical test on outputs alone can expose a well-trained generator, provenance must be embedded at generation time. We introduce Tabular Pancakes, a training-free watermark that replaces the generator's Gaussian latent prior with the homogeneous CLWE distribution: Gaussian in all directions except a secret subspace, where it is periodic. With the key, an auditor efficiently recovers the watermark from a small fraction of released rows, and every verdict carries a certified p-value. Recovery is invariant to row permutation, column renaming, and rescaling. Under tampering such as cell edits or re-imputation, detection power decreases only linearly in the fraction modified—no small edit removes the mark. Row-level tests localize synthetic rows in mixed tables with calibrated false-discovery rates, turning detection into forensic evidence. The mark can equivalently be embedded directly in released rows, covering generators whose samplers cannot be inverted. Without the key, marked and unmarked outputs are indistinguishable, and keyless removal provably requires √n/k-fold larger distortion than keyed removal, ruling out selective scrubbing. We verify Tabular Pancakes on four datasets and six generator families (TabSyn, TabDDPM, STaSy, CoDi, SMOTE, GReaT), towards making data fabrication provably traceable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.