acceptodds
Under review as a conference paper at ICLR 2027

Joint, Decoupled, or Agent: When Each Multimodal Alignment Paradigm Wins

Abstract

Jointly pretrained multimodal encoders are a common way to combine modalities, while alternatives integrate their information during downstream training or inference. When does each approach win? We compare three paradigms, instantiated with public checkpoints, across 16 paired datasets in five domains and multiple training-label budgets. Joint trains a fusion head on frozen, jointly pretrained encoders. Decoupled trains a fusion head on frozen, independently pretrained encoders. Agent gives a frozen LLM task instructions, labelled examples, available sample text, and supervised unimodal readouts. No paradigm dominates the complete dataset–budget combinations: Decoupled wins , Joint , and Agent by mean utility. ScienceQA, MELD, and MetMeme shift from Agent to fusion as supervision increases. Between the fusion paradigms, Joint leads on four of five QA datasets across all evaluated budgets, while higher unimodal -NN accuracy is associated with a relative shift toward Decoupled across eight classification datasets (Spearman ). We test whether representation and task diagnostics predict these preferences on held-out datasets before fitting their candidate systems. A two-step cascade combines these diagnostics with task and budget metadata to select Agent or fusion, then Joint or Decoupled. Under leave-one-group-out evaluation, it selects the winner in of cells versus for metadata-only routing, with mean relative regret versus . Using geometry, supervised representation diagnostics, and text–label matching together yields higher accuracy and lower regret than using any one of these sources with metadata. These results connect paradigm preferences to supervision budget, task format, and unimodal -NN accuracy. Target-task diagnostics thus support selecting among the three paradigms before candidate fitting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.