Flow-Embedded Multimodal Variational Autoencoders for Incomplete Multi-view Clustering
Abstract
Incomplete multi-view clustering aims to uncover the underlying cluster structure from data observed across multiple views with missing modalities. Recently, generative modeling has emerged as a promising paradigm for this task, as it provides a principled framework for capturing complex data distributions and inferring unobserved information from the available views. In this paper, we propose **F**low-**E**mbedded **M**ultimodal **VA**riational auto**E**ncoders (FEMVAE) with Gaussian mixture priors to learn expressive latent distributions and a coherent shared clustering structure from incomplete multi-view data. Specifically, we integrate normalizing flows into variational autoencoders to construct flexible posterior distributions within a unified framework, while leveraging a Gaussian mixture prior to impose a cluster-aware structure on the learned latent space. Moreover, to fully align the view-specific latent representations, we employ a pairwise optimal transport strategy to promote distributional consistency across views. This enables the complementary information from different views to be effectively integrated despite missing modalities. Once trained, our model can naturally recover missing data modalities. Extensive experiments on several benchmark incomplete multi-view datasets demonstrate the effectiveness of the proposed method, which consistently achieves competitive clustering performance against state-of-the-art approaches under different missing-view settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.