acceptodds
Under review as a conference paper at ICLR 2027

Forced Fusion: Multimodal Models Do Not Know When Not to Bind

Abstract

Multimodal models benefit from complementary observations, but unrelated evidence can redirect predictions away from the intended target. We call this failure Forced Fusion. We demonstrate it in seven task-specific fusion architectures on a synthetic task and five audio-visual datasets, and in pretrained general-purpose multimodal models on audio-visual and audio-text tasks. Our analysis of matched training motivates learning source correspondence to determine when auxiliary evidence should influence prediction. We propose a module based on Bayesian causal inference (BCI) that explicitly learns correspondence and uses it to balance fused and target-only predictions. Evaluations show that our approach mitigates Forced Fusion and improves average target accuracy over the implicit Mixed baseline in most audio-visual settings. Furthermore, analyses of task-specific and adapted general-purpose models show that explicit supervision helps models distinguish matched from mismatched inputs, even when gains in target accuracy are modest. Controlled interventions keep the module's branch predictions fixed and vary only its mixing weight, showing that correspondence-dependent weighting improves mismatched prediction on average. These findings highlight the importance of evaluating target accuracy alongside the ability to recognize source correspondence and use it to guide multimodal fusion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.