Decoupling Is Not Enough: Understanding Generalization in Multimodal Fusion
Abstract
Modality imbalance often prevents multimodal learning from fully exploiting the information available in each modality. Although decoupling unimodal learning from fusion alleviates optimization interference, it does not ensure fusion generalization. We identify a remaining source of failure: fusion trained on samples used for unimodal fitting can exploit modality reliability and error patterns that transfer poorly to unseen data. We analyze this failure mechanism and propose a generalization-aware multimodal fusion framework. The framework constructs out-of-fold representations and predictions to learn feature-based corrections to averaged unimodal logits from out-of-sample evidence. To account for differences among the unimodal models fitted across folds, we formulate a cross-fold objective that balances average correction gains against unfavorable fold-specific effects. Experiments across multimodal benchmarks support the effectiveness of the framework, with component ablations demonstrating the complementary benefits of out-of-fold fusion learning and cross-fold regularization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.