LAFE: Layer Augmented Fusion Estimation
Abstract
Multimodal learning frequently suffers from modality imbalance, which existing methods primarily study during model training. However, we find that modality imbalance may prevent final predictions from fully exploiting the class information already learned in modality representations. We observe that predictions from some intermediate layers correctly classify examples that no convex combination of the original unimodal and fusion head outputs can classify correctly. Nevertheless, we show that high confidence training fits attenuate training loss responses to these predictions under bounded score changes. These observations suggest that fully exploiting learned modality information requires considering both which predictions to fuse and how to estimate their combination. We propose Layer Augmented Multimodal Fusion (LAMF), which fits linear probes to intermediate representations of frozen encoders to complement the original classifier outputs. LAMF applies base models and their probes to samples excluded from their fitting and uses the resulting predictions to estimate a global combination of all prediction sources. The method preserves the base architecture and training objective, and inference uses only one full model and its probes. Extensive experiments demonstrate that LAMF outperforms current state-of-the-art methods on multiple tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.