Imperfect Alignment, Reliable Similarity: Score–Encoder Separation in Multimodal Contrastive Learning
Abstract
Multimodal contrastive learning is often interpreted as aligning matched representations, yet each modality may contain noise irrelevant to cross-modal comparison. We ask whether reliable similarity requires the individual encoders to be denoised. We study population gradient flow for linear encoders with a learned temperature in a noisy two-modality model, where the score is factorized through the encoders. In a regular non-collapsing regime, every nuisance interaction in the score vanishes asymptotically while shared score signal remains. Generic initialization keeps encoder nuisance nonzero at every finite time, but permits asymptotic encoder denoising. Separately, when the output dimension exceeds the shared-signal dimension by at least two, an open basin of full-signal optima retains nuisance in both encoders asymptotically. Thus score- and encoder-level denoising are distinct. Controlled linear and nonlinear experiments support this separation; a real-data compression study further separates feature reconstruction, preservation of learned comparisons, and label retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.