E-MIND: Conditional Energy Matching for Multimodal Imputation and Denoising
Abstract
Missing and corrupted modalities pose a common challenge for multimodal sentiment analysis: recovering task-relevant features from the remaining modalities. They nevertheless provide different starting points, as a missing modality offers no target observation, whereas a corrupted modality retains an imperfect one. We introduce E-MIND, a unified recovery framework that adapts Energy Matching to cross-modal conditional imputation and denoising. For each target modality, E-MIND jointly learns relative feature energies and recovery directions: contrastive energy learning encourages lower energies for real features than sampled negatives, while conditional flow regression supervises the negative energy gradient along noise-to-data interpolations. Conditioned on the remaining clean modalities, the resulting energy field generates missing features from Gaussian noise and refines corrupted features from their observations. Both modes reuse the same modality-specific energy parameters with different sampling settings and require no separate noise-specific training. Recovered and retained features are then fused for sentiment prediction. On CMU-MOSI, E-MIND achieves the highest weighted F1 among the compared methods at all seven nonzero nominal missing rates; on CMU-MOSEI, it achieves the highest seven-class accuracy in seven of eight settings. A recovery-loss ablation shows that combining the two objectives outperforms either alone, while separate single- and paired-modality corruption evaluations support reusing the learned fields for observation-based refinement. An extension of E-MIND to AVMNIST outperforms zero-fill and mean-fill baselines in classification accuracy, demonstrating the applicability of conditional energy recovery beyond sentiment analysis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.