acceptodds
Under review as a conference paper at ICLR 2027

Alignment Strength Is Not Enough: Backward Influence Allocation in Heterogeneous Multimodal Learning

Abstract

Multimodal alignment is commonly controlled through a scalar strength, yet paired modalities need not contribute equally useful information. This paper studies a complementary control dimension: backward influence allocation. We introduce encoder-specific admission variables that determine how strongly a shared alignment gradient enters each encoder while leaving the forward alignment objective unchanged. We show that isotropic admission is equivalent to scalar tuning, whereas asymmetric admission accesses encoder updates that a single scalar weight cannot in general reproduce. Dense multi-seed comparisons reveal residual representation trade-offs under heterogeneous sensing, while matched sensing can be nearly scalar-sufficient, indicating that the value of allocation is condition dependent. Motivated by this observation, we develop a lightweight training-time policy that maps task–alignment gradient interaction to encoder-specific admission levels and learns from training-only progress, without preassigning a teacher modality or adding inference-time components. On ViT-B/16 across RGB–event and RGB–sonar retrieval, the learned controller attains the highest mean validation and ID R@5 among five specialized multimodal optimization baselines on all three datasets. Against validation-tuned scalar alignment, it achieves higher mean ID retrieval on all three datasets and higher mean validation retrieval on two. These results establish backward influence allocation as a distinct and learnable control mechanism for heterogeneous multimodal alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.