STR: Auditing Shared-Direction Geometry in Multimodal Alignment
Abstract
Multimodal contrastive learning involves inter-modal semantic agreement, inter-modal separation, and intra-modal diversity, yet most alignment objectives fold these concerns into a single aggregate similarity score. This paper decomposes them into separately weighted geometric terms, collectively Surface-Tension Regularization (STR), an auxiliary regularizer with no additional learnable parameters, and uses the decomposition as a measurement instrument. Across four VAST benchmarks and three seeds per benchmark, canonical STR increases common-mode energy and differs from the baseline in its response to shared-direction removal. Separating deletion from residual-norm rescaling in these two models shows that deletion drives the baseline loss, whereas under STR rescaling alone is disruptive and the combined operation approximately preserves aggregate top-1 recall. Replacing Gram-volume attraction with pairwise cosine also yields energy inflation and a similar combined-removal response on MSR-VTT and DiDeMo across three seeds per dataset. Seed-0 factorial and checkpoint-conditional gradient measurements provide complementary evidence about coupling and mean-margin changes. A single fixed STR configuration without per-dataset tuning produces small changes in final top-1 recall on the four benchmarks. The same metric remains similar across configurations with substantially different embedding-space behavior, motivating separate evaluation of the embeddings deployed directly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.