Robust Multimodal Learning via Dual-Constrained Informative Latent Space
Abstract
Multimodal systems deployed in real-world scenarios often encounter both missing modalities and corrupted observations. Such inputs make it challenging for existing methods to identify reliable information for effective multimodal interaction. Moreover, noisy and missing modalities are typically handled separately, hindering unified robustness under complex multimodal conditions. To address these challenges, we propose Reliability-aware Dual-constraint Information Bottleneck (RDIB). RDIB introduces a dual-constraint information bottleneck that filters noisy information while preserving both target-modality and task-relevant information, providing an information-theoretic basis for representation reliability. Building on the informative representations, a unified reliability estimation mechanism combines bottleneck-derived reliability for observed representations with estimated reliability for reconstructed representations. We further construct reliability-aware complete and incomplete fusion streams and impose a fusion-level dual-constraint information bottleneck to learn robust multimodal representations under incomplete inputs. Extensive experiments on MOSEI, D-Vlog, LMVD and Food-101 under diverse modality conditions demonstrate that RDIB consistently achieves robust and competitive performance across multimodal learning scenarios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.