acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Modality-Aware Dual-Noise Training for Robust Multimodal Sentiment Analysis

Abstract

Multimodal Sentiment Analysis (MSA) aims to infer human emotions by integrating text, audio, and visual modalities. However, real-world deployment remains challenging due to heterogeneous perturbations, such as random data loss and transmission-induced feature distortion in distributed sensing environments. Existing methods frequently presuppose static modality dominance and focus on mitigating isolated forms of noise. To address the diverse and complex noise perturbations encountered in real-world scenarios, we propose a Dynamic Modality-Aware Dual-Noise Training framework (DMDN). Our framework first introduces the Dynamic Information-guided Modality Correction Module (DIMC), which estimates modality degradation via KL divergence and dynamically identifies the most reliable modality to enhance corrupted features. In this framework, modality features are encoded into sparse spike trains by a spiking neuron-based sender, transmitted through simulated noisy channels, and decoded by a State Space Model (SSM)-based receiver. Our framework jointly models random data missing and channel noise, better reflecting the complex perturbation patterns in real-world distributed systems. Empirically, joint training under such multi-noise conditions encourages the model to suppress irrelevant perturbations while preserving task-relevant information. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art baselines across multiple MSA benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.