acceptodds
Under review as a conference paper at ICLR 2027

Robust multi-modal models by explicitly optimizing uniformity and alignment

Abstract

We define the robustness of a multi-modal embedding using the margin of one nearest neighbor retrieval and derive the form of the optimal robust embedding (ORE) for various finite dataset sizes () and dimensions (). Specifically, we show that for the most robust embedding is one where the modalities are aligned and the embeddings form the vertex of a regular simplex. Previous work has shown that the ORE is also the global optimum of the CLIP loss yet gradient-based optimization of the CLIP loss does not converge to the ORE, and instead learns embeddings with a modality gap. To alleviate this failure, we analyze an alternative loss, Explicit Multi-Modal Uniformity and Alignment (EMMUA), whose global optimum is also the ORE. We show both theoretically and experimentally that gradient-based optimization on the EMMUA does not learn a modality gap and finds significantly more robust embeddings than those obtained with the CLIP loss.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.