acceptodds
Under review as a conference paper at ICLR 2027

Unpaired Multimodal Learning with Cross-Modal Self-Supervision

Abstract

Multimodal learning typically relies on paired observations, yet many domains, including biology and medicine, provide abundant unimodal data with only sparse and potentially imperfect pairing. We introduce unimo-x, a general learning framework for training multimodal Transformers on highly unpaired corpora. Our approach uses teacher-anchored cross-modal self-supervision to turn unpaired observations into training signals, complementing paired supervision while mitigating steganographic shortcuts that can undermine consistency-based learning. We instantiate unimo-x in 4M and evaluate it across several datasets and tasks. Unpaired observations improve image generation and caption semantics over paired-only training.With 25%, 17.5%, or 10% paired samples, unimo-x recovers substantial benefits of paired supervision, from competitive image-generation quality with four times fewer pairs to stronger caption representations than paired-only training even at 10% pairing. These findings point to a practical route to multimodal learning in which sparse paired examples guide training on much larger unimodal corpora.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.