Robust Multimodal Representation Learning through Topology-Preserving Prototype Maps
Abstract
Learning representations that preserve relations between samples is a fundamental challenge in machine learning. For multimodal data, the representations should retain the information of each modality while remaining aligned across modalities. A common approach is to map all inputs into a shared latent space. However, this requires them to share a common geometry and may lead to the loss of modality-specific information. In this work, we propose Pesto, a representation learning method that uses a shared topology for alignment while learning a separate embedding space for each modality. We represent the topology as a grid of prototypes, where similar samples are mapped to neighbouring locations. For multimodal data, each modality has its own encoder and prototype matrix, and modalities are aligned through their distributions over the shared map. Thus, different modalities of the same instance are aligned to similar locations on the map, while similar samples are organised into neighbouring regions. We learn the representations and topology jointly by backpropagating the neighbourhood loss through the encoders and using it to update the prototypes. Since alignment takes place on the shared map rather than embedding spaces, each modality can be used separately at inference time without imputation or retraining. Experiments on real-world datasets show that our method retains the characteristics of individual modalities, produces consistent topological maps across modalities, and achieves competitive performance compared with existing methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.