Evidential Multi-view Uncertainty Learning for Multimodal Fake News Detection
Abstract
Multimodal fake news detection benefits from complementary textual and visual cues; however, existing methods often treat local text tokens and image patches as equally reliable. Consequently, intra-modal noise may propagate through cross-modal interaction and impair final decisions. To overcome this challenge, we present Evidential Multi-view Uncertainty Learning (EMUL), a unified framework that seamlessly integrates uncertainty as a continuous regulatory signal across representation learning, semantic alignment, and decision fusion. Rather than processing components in isolation, EMUL first estimates token- and patch-level intra-modal uncertainty to perform confidence-weighted aggregation for fine-grained noise suppression. These estimated confidences directly condition cross-modal semantic alignment, reducing the distribution gap between reliable local modality representations and global image-text semantics while down-weighting uncertain samples. Ultimately, at the decision stage, this uncertainty-awareness enables EMUL to jointly calibrate the text, image, and fused views using intra-modal and evidential uncertainty prior to multi-view evidence fusion. Extensive experiments on three public benchmarks demonstrate that EMUL achieves state-of-the-art performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.