acceptodds
Under review as a conference paper at ICLR 2027

Channel-Aware Turbo Multimodal Fusion: Learning under Noisy and Missing Modalities

Abstract

Multimodal systems in the wild must operate when modalities are corrupted by noise or entirely missing. Most existing methods address the two failure modes separately and fuse modalities with attention that is blind to signal quality. We draw an analogy between multimodal fusion and multi-channel communication, and propose Channel-Aware Turbo Multimodal Fusion (CAT-MF), which treats noise and missingness uniformly as channel impairments. CAT-MF has three components: (1) channel state estimation, which scores every modality by an existence probability, a spectral proxy for the inverse signal-to-noise ratio, and a learned damage-type embedding, all obtained without clean reference signals; (2) iterative joint decoding, a Turbo-inspired procedure that exchanges soft cross-modal information over a small number of rounds and accumulates it in a persistent state; and (3) channel-modulated attention, which gates cross-modal transfer by the estimated states, zeroing missing channels, down-weighting noisy ones, and suppressing transfer between channels sharing correlated damage. We prove that the decoding iterations converge absolutely and linearly to a limit once the steps are summable, that the accumulated extrinsic evidence yields non-negative information gain per round bounded below by the donor modalities' quality, and that the reconstruction error decays with the availability of high-quality correlated modalities. CAT-MF exceeds competitive baselines on multiple datasets and is markedly more robust under noisy and missing modalities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.