Leveraging Generative AI for Robust Clinical Predictions: A Case Study on Dental Implant Failures
Abstract
Machine learning has the potential to improve clinical prediction and decision- making, yet its real-world effectiveness is often constrained by the characteris- tics of healthcare data. Clinical datasets are frequently small, heterogeneous, and highly imbalanced, limiting model generalization and reducing the detection of rare but clinically significant outcomes. These challenges motivate the develop- ment of learning frameworks that can effectively leverage limited data while pre- serving the statistical characteristics necessary for robust prediction. This paper investigates synthetic tabular data generation as a strategy for improving predic- tive modeling under data scarcity and class imbalance. We propose GC-TabGAN (Gaussian-Conditional Tabular GAN), a customized Generative Adversarial Net- work (GAN) for generating realistic synthetic clinical records, and integrate it with a Feed Forward Neural Network (FFNNet Baseline) for downstream pre- diction. Dental implant failure is used as a representative clinical case study be- cause implant failures are relatively rare, arise from complex interactions among patient, clinical, and procedural factors, and therefore present a challenging pre- diction task. Using a public dataset of 747 implants, of which 117 (15.7%) failed, and a 10-fold cross-validation protocol in which the generator is refit within every training fold, we add 5%, 25%, 50%, and 90% of a 1,000-record fold-specific syn- thetic pool to the real training data. Synthetic augmentation mainly improves the detection of implant failures. Using only real data, the FFNNet Baseline achieved an accuracy of 0.92, an Area Under the Receiver Operating Characteristic Curve (AUC) of 0.81, and a sensitivity to implant failure of 0.64. Augmenting the train- ing data with synthetic samples raised the sensitivity to 0.76–0.80 while maintain- ing a specificity of 0.97–0.98; with 50% synthetic samples, the model reached an accuracy of 0.95, an AUC of 0.89. In contrast, oversampling the training data with SMOTE raised sensitivity to 0.88 but reduced specificity to 0.66, with an AUC of 0.77, well below that of GC-TabGAN, indicating that GC-TabGAN improves dis- crimination rather than merely shifting the decision boundary toward the minority class. These findings suggest that distribution-aware synthetic tabular data can partly offset data scarcity and class imbalance in small clinical datasets, a setting common in healthcare where large, balanced datasets are difficult to collect.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.