KASS: -Anonymity-Aware Sequential Sampling for Synthetic Tabular Data
Abstract
Synthetic tabular data generation seeks to preserve key statistical properties of the original data while reducing disclosure risk, yet existing generators may reproduce rare or unique attribute combinations and violate hard structural constraints. We propose KASS, a -anonymity-aware sequential sampling framework for synthetic tabular data. Rather than enforcing -anonymity as a post-processing step, KASS embeds its protective logic directly into the sampling process by restricting candidate values to attribute combinations with sufficient empirical support, making rare and potentially identifying patterns structurally unreachable. To preserve cross-variable dependencies and data plausibility, KASS employs normalized mutual information (NMI) thresholding to relax conditioning variables when strict conditioning becomes too sparse, and supports multiple banned-set strategies for controlling the privacy-utility trade-off. We analyze KASS from a differential privacy perspective and establish its relationship with -diversity. Experiments on real-world datasets show that KASS achieves competitive utility and distributional fidelity, produces substantially fewer constraint violations than competing methods, and generates no unique original samples.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.