acceptodds
Under review as a conference paper at ICLR 2027

DIDA: A Value-Driven Paradigm for Data Importance-Based Tabular Data Augmentation in Low-Data High-Constraint Scenarios

Abstract

In regulated domains such as medical diagnosis and financial risk control, tabular datasets are often small, imbalanced, and subject to strong cross-feature constraints. Existing augmentation methods therefore face a difficult trade-off among preserving semantic consistency, increasing sample diversity, and keeping augmentation decisions auditable. We propose DIDA (Data-Importance driven Differential Augmentation), a value-aware framework that differentiates augmentation at the sample-feature level using two complementary signals: feature-level SHAP importance and sample-level Shapley contribution. We formulate augmentation as diversity maximization under semantic-consistency constraints and show that the optimal perturbation magnitude decreases with sample contribution and scales inversely with feature importance (Theorem 1), providing a principled basis for value-dependent perturbation. DIDA further incorporates dynamic domain correction and constraint-guided categorical augmentation to preserve valid feature combinations. Across six public tabular datasets, six augmentation baselines, and ten downstream classifiers, DIDA improves average AUC-ROC, AUC-PR, and F1 by 1.94, 2.30, and 4.30 percentage points, respectively, over the strongest competing baseline. Its generated samples also achieve lower JS divergence and Wasserstein distance, together with higher information entropy and feature coverage, indicating that DIDA improves diversity without sacrificing distributional fidelity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.