TabMaDo: Dual-Guided Tabulamba Diffusion Oversampling for Imbalanced Tabular Data Classification
Abstract
Tabular data, as a prevalent form of structured data, often suffers from class imbalance that limits classification performance. To address this issue, diffusion models have been employed for oversampling due to their ability to capture complex feature dependencies. However, when minority class samples are extremely scarce, they struggle to learn the genuine distribution of the minority class, consequently yielding generated samples of poor quality and diversity. To mitigate this challenge, this study proposes TabMaDo, a dual-guided diffusion framework tailored for tabular data. The framework consists of two components, a Mamba-based noise predictor named Tabulamba incorporates Class-Aware Mamba Blocks that dynamically highlights minority class information, preventing the generation process from being overwhelmed by the majority class. This enables the diffusion model to learn the overall data distribution, moving beyond the constraints imposed by scarce minority class samples. Complementing this, a lightweight Guider is introduced to provide dual guidance on both class label and distribution. This ensures the reverse diffusion process is steered toward minority class characteristics while preserving sample diversity, effectively overcoming the limitation of minority class sample scarcity. Together, they form a cohesive framework that enables high-quality sample generation under extreme class imbalance. Extensive experiments on 10 binary classification datasets demonstrate that TabMaDo outperforms six SOTA baseline methods across F1-score, G-mean and AUC metrics, achieving maximum improvements of 12% in F1-score, 16.5% in G-mean and 0.45% in AUC. These results confirm that TabMaDo provides an effective solution for imbalanced tabular data generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.