Tabular-to-Graph Diffusion: Structured Generative Oversampling for Imbalanced Tabular Data
Abstract
Class imbalance presents severe challenges for supervised classification on tabular data. Traditional oversampling techniques synthesize minority records via linear interpolation, which often crosses decision boundaries in non-linear distributions. Tabular diffusion models alleviate this by learning generative densities; however, flattening records into one-dimensional vectors and initiating reverse trajectories from standard Gaussian noise across the full feature space frequently leads to uncalibrated score estimates in low-density regions, resulting in mode drift toward majority classes and damaged column correlations. We propose Tabular-to-Graph Diffusion (TGD), an oversampling framework designed to preserve feature dependencies under acute sample scarcity. TGD structures each minority record alongside its nearest minority neighbors into a local feature graph, where nodes represent tabular columns and node attribute vectors contain feature values across ordered neighboring records. A two-layer residual graph convolutional network updates feature representations using empirical correlation coefficients and a categorical mask that isolates mutually exclusive dummy variables, while linear layers project across neighbor coordinates to capture local variations. Reverse diffusion begins from a localized perturbation of the empirical neighborhood and integrates along deterministic trajectory paths, preventing synthetic records from drifting outside populated minority regions. Evaluations across eight benchmark datasets show that TGD achieves competitive downstream classification performance across heuristic interpolators, deep generative models, and vector-based tabular diffusion methods, establishing a practical trade-off between decision boundary adherence and multi-feature correlation preservation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.