Unpaired Multimodal Learning with Cross-Modal Self-Supervision
Abstract
Multimodal learning typically relies on paired observations, yet many domains, including biology and medicine, provide abundant unimodal data with only sparse and potentially imperfect pairing. We introduce unimo-x, a general learning framework for training multimodal Transformers on highly unpaired corpora. Our approach uses teacher-anchored cross-modal self-supervision to turn unpaired observations into training signals, complementing paired supervision while mitigating steganographic shortcuts that can undermine consistency-based learning. We instantiate unimo-x in 4M and evaluate it across several datasets and tasks. Unpaired observations improve image generation and caption semantics over paired-only training.With 25%, 17.5%, or 10% paired samples, unimo-x recovers substantial benefits of paired supervision, from competitive image-generation quality with four times fewer pairs to stronger caption representations than paired-only training even at 10% pairing. These findings point to a practical route to multimodal learning in which sparse paired examples guide training on much larger unimodal corpora.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.