COIN: Coordinating Multimodal Updates through Functional Invariance
Abstract
Multimodal learning must not only train each modality well, but also coordinate how their updates affect the shared prediction. Joint training can underuse informative modalities, while improving each modality separately does not ensure that their updates work well together after fusion. We show that module-wise signals can be misleading: coordinate rescaling can create arbitrarily different learning speeds, and large module responses can cancel while producing little or no fused-output change. To study this interaction directly, we characterize each module by the first-order change it can induce in the fused output and identify a *functional invariant relation*: different module responses can produce the same combined output change. Among such equivalent responses, the minimum-energy allocation removes unnecessary module movement without changing the prediction to first order. Building on this principle, we derive **COIN**, which jointly chooses a task-improving output change and how to distribute it across modules. COIN accounts for both each module's relation to the task and its interaction with other modules through a small linear system. This view also explains why a module that cannot improve the task on its own may still help by correcting another module's response. We further bound total module response and derive the exact cost of ignoring cross-module interactions. Across evaluated benchmarks, COIN achieves the highest mean fusion accuracy among the compared methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.