acceptodds
Under review as a conference paper at ICLR 2027

Preserving Unique Concepts in Different Modalities for Contrastive Learning: Improvements to Contrastive Audio-Visual Masked Autoencoder

Abstract

Contrastive learning is widely adopted in multimodal model training to narrow the distance between cross-modal feature representations, facilitating multimodal information fusion and unified representation learning. During training, contrastive loss guides the model to capture matched features of the same concept across different modalities. Nevertheless, it also causes the model to discard modality-specific conceptual information that exists solely within a single modality and cannot be shared cross-modally. To address this limitation, we decouple each modality’s features into shared conceptual components and modality-exclusive components. Cross-modal contrastive alignment is only imposed on shared components, which preserves modality-specific semantics while sustaining the model’s cross-modal reasoning capacity. Our experiments verify that globally aligning full features via contrastive loss degrades downstream task accuracy, whereas retaining modality-unique concepts yields steady performance improvements. Built upon the audio-visual multimodal model CAVMAE, we redesign its contrastive learning module to preserve exclusive conceptual information from both visual and audio modalities. Experiments across multiple datasets demonstrate that our method outperforms the vanilla baseline on downstream task. The results validate the rationality and effectiveness of our approach for audio-visual multimodal learning. Further analyses confirm that audio and visual modalities contain distinct exclusive concepts, which substantially boost the model’s general representation capability and downstream performance. Overall, our work proves the efficacy of preserving modality-specific conceptual information, especially for audio-visual multimodal tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.