acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Text Semantics in Supervised Unpaired Multimodal Representation Learning

Abstract

Supervised *Unpaired Multimodal Representation Learning* (UMRL) improves visual classification using text representations, requiring neither *prior cross-modal alignment* nor *sample-level correspondence*. Under such conditions, what drives the performance gains remains unclear, particularly the role of text semantics. In this paper, we revisit supervised UMRL and identify visual representation adaptation and text-derived classifier initialization as the main contributors, while increasing the number of text representations and introducing a text classification loss provide only limited benefits. Interestingly, shuffling the correspondence between text representations and class labels has little effect on the text-derived classifier initialization. Motivated by this, we introduce a semantics-agnostic representation synthesis strategy for classifier initialization. Extensive experiments show that semantics-agnostic initialization performs comparably to its text-derived counterpart. It also eliminates the need for an additional text encoder and reduces representation construction time and storage overhead. Our observations suggest that text semantics may not be the primary driver of the performance gains in supervised UMRL. The codes are available at https://anonymous.4open.science/r/RethinkUMRL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.