Unimodal Augmentation-based SSL Objectives Induce Cross-Modal Alignment
Abstract
Recent works have observed implicit cross-modal alignment between independently trained unimodal models. However, the mechanisms underlying the emergence of this phenomenon remain unresolved. In this paper, we demonstrate that unimodal augmentation-based self-supervised objectives implicitly induce cross-modal alignment. We show that low unimodal losses force high cross-modal alignment, whereas a high unimodal loss prevents such alignment. We theoretically connect the widely used InfoNCE loss to local CKA, the popular measure of cross-modal alignment. We further provide extensive empirical evidence supporting this claim, both in controlled settings and across 53 unconstrained frozen text, vision, and audio unimodal encoders. In both cases, unimodal augmentation-based objectives control cross-modal alignment. Overall, our results provide a theoretical and empirical foundation for understanding the emergence of implicit cross-modal alignment by establishing a direct connection between unimodal training objectives and cross-modal structure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.