acceptodds
Under review as a conference paper at ICLR 2027

STR: Auditing Shared-Direction Geometry in Multimodal Alignment

Abstract

Multimodal contrastive learning involves inter-modal semantic agreement, inter-modal separation, and intra-modal diversity, yet most alignment objectives fold these concerns into a single aggregate similarity score. This paper decomposes them into separately weighted geometric terms, collectively Surface-Tension Regularization (STR), an auxiliary regularizer with no additional learnable parameters, and uses the decomposition as a measurement instrument. Across four VAST benchmarks and three seeds per benchmark, canonical STR increases common-mode energy and differs from the baseline in its response to shared-direction removal. Separating deletion from residual-norm rescaling in these two models shows that deletion drives the baseline loss, whereas under STR rescaling alone is disruptive and the combined operation approximately preserves aggregate top-1 recall. Replacing Gram-volume attraction with pairwise cosine also yields energy inflation and a similar combined-removal response on MSR-VTT and DiDeMo across three seeds per dataset. Seed-0 factorial and checkpoint-conditional gradient measurements provide complementary evidence about coupling and mean-margin changes. A single fixed STR configuration without per-dataset tuning produces small changes in final top-1 recall on the four benchmarks. The same metric remains similar across configurations with substantially different embedding-space behavior, motivating separate evaluation of the embeddings deployed directly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.