Rethinking Text Semantics in Supervised Unpaired Multimodal Representation Learning
Abstract
Supervised *Unpaired Multimodal Representation Learning* (UMRL) improves visual classification using text representations, requiring neither *prior cross-modal alignment* nor *sample-level correspondence*. Under such conditions, what drives the performance gains remains unclear, particularly the role of text semantics. In this paper, we revisit supervised UMRL and identify visual representation adaptation and text-derived classifier initialization as the main contributors, while increasing the number of text representations and introducing a text classification loss provide only limited benefits. Interestingly, shuffling the correspondence between text representations and class labels has little effect on the text-derived classifier initialization. Motivated by this, we introduce a semantics-agnostic representation synthesis strategy for classifier initialization. Extensive experiments show that semantics-agnostic initialization performs comparably to its text-derived counterpart. It also eliminates the need for an additional text encoder and reduces representation construction time and storage overhead. Our observations suggest that text semantics may not be the primary driver of the performance gains in supervised UMRL. The codes are available at https://anonymous.4open.science/r/RethinkUMRL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.