Efficient Multimodal Fusion for Data-Constrained Regimes
Abstract
Multimodal classification often requires substantial paired data to learn how to combine heterogeneous representations from different modalities. We consider a streamlined regime in which modality-specific encoders are frozen and adaptation operates directly within their pretrained representation spaces. Our method, Factorized Modality Evidence targets multimodal fusion in data-constrained settings through a factorized multimodality approach. We model class structure independently within each modality and combines the resulting evidence in the shared label space. Each modality contributes a Gaussian discriminant score, while a single scalar coefficient calibrates the modality evidence and class prior. This avoids representation-level alignment, encoder updates, and high-dimensional learned fusion, while naturally supporting missing modalities by omitting unavailable evidence terms. Across six datasets and eleven classification tasks, FaME outperforms the strongest competing baseline in 36 of 42 few-shot dataset–shot settings. Under test-time modality missingness, our method consistently outperforms standard imputation and modality-dropout baselines. When few-shot adaptation data contain missing modalities, FaME achieves the strongest performance across all 18 settings. These results show that lightweight factorized adaptation can effectively exploit complementary information from independently pretrained modalities without learning a high-capacity multimodal fusion model. Our code is available at: https://anonymous.4open.science/r/FaME-66F3/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.