Integrated PU Learning and Deep Generative VAEs for Indication-Driven Missingness to Improve MNAR Imputation and Prediction
Abstract
Recent advances in deep generative modeling have improved imputation for missing not at random (MNAR) data. However, jointly improving both imputation quality and downstream prediction remains challenging. In this paper, we propose to use the positive-unlabeled (PU) learning to deal with indication-driven missingness to leverage partially observed measurement-intent variable to improve MNAR imputation and downstream prediction. We first use a PU adversarial network to learn probabilities of unknown positives of intention to measure among all the missing cases, which then guide the training of variational autoencoders (VAE) that jointly model observed data and missingness indicators. Experiments on synthetic and real-world data show that our proposed framework has superior performance for both imputation and downstream prediction tasks. These results show that explicitly modeling missingness indicators integrated with additional PU measurement intent information can improve MNAR imputation and downstream predictions under various scenarios and assumptions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.