Robust multi-modal models by explicitly optimizing uniformity and alignment
Abstract
We define the robustness of a multi-modal embedding using the margin of one nearest neighbor retrieval and derive the form of the optimal robust embedding (ORE) for various finite dataset sizes () and dimensions (). Specifically, we show that for the most robust embedding is one where the modalities are aligned and the embeddings form the vertex of a regular simplex. Previous work has shown that the ORE is also the global optimum of the CLIP loss yet gradient-based optimization of the CLIP loss does not converge to the ORE, and instead learns embeddings with a modality gap. To alleviate this failure, we analyze an alternative loss, Explicit Multi-Modal Uniformity and Alignment (EMMUA), whose global optimum is also the ORE. We show both theoretically and experimentally that gradient-based optimization on the EMMUA does not learn a modality gap and finds significantly more robust embeddings than those obtained with the CLIP loss.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.