SMOOTHALIGN: PROBABILISTIC SMOOTHING FOR MULTIMODAL REPRESENTATION LEARNING AND ALIGNMENT
Abstract
Multimodal representation learning and alignment are fundamental to cross-modal understanding and retrieval. However, common multimodal alignment objectives represent samples as deterministic points and assign every unpaired observation the same negative status: semantically related but identity-different samples are repelled as strongly as unrelated ones. We identify this structural failure as the semantic overlap bias. To correct this bias, we propose a probabilistic method, SmoothAlign, which uses a confidence-calibrated distributional prior derived from Gaussian semantic clouds of intermediate encoder representations to assign different repulsion strengths to unpaired samples according to their semantic overlap. We further prove that SmoothAlign admits a conservative mutual-information lower bound. SmoothAlign consistently outperforms the corresponding baselines in both video retrieval and image–text retrieval. A controlled IEMOCAP diagnostic confirms that the correction assigns stronger softening to semantically overlapping non-pairs than to ordinary negatives. The code is available at https://anonymous.4open.science/r/SmoothAlign-4EF0/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.