EvidenceSAM: Hierarchical Source Evidence Fusion with SAM for Multimodal Semantic Segmentation
Abstract
Multimodal semantic segmentation improves dense scene understanding by combining complementary observations from heterogeneous sensors. However, the usefulness of different modalities is inherently spatially and semantically dependent: a source that is informative for one region or category may be redundant or unreliable for another. Most existing approaches address this challenge primarily through feature-level interaction, where source-specific information is progressively compressed into a shared representation. This becomes particularly important when adapting large segmentation foundation models, where multimodal reasoning should exploit complementary sensing cues without sacrificing the transferable representations learned during pretraining. We introduce EviSAM2, a parameter-efficient SAM2-based framework that preserves source-specific evidence beyond the initial feature-fusion stage. Rather than treating multimodal fusion as a single operation, EviSAM2 organizes prediction into three complementary levels: Cross-Modal Feature Fusion (CMF) constructs the shared multimodal representation, Risk-Decomposed Expert Aggregation (RDEA) integrates source-specific semantic evidence, and Context-Guided Semantic Refinement (CGSR) further refines the aggregated prediction using class-level context. This formulation explicitly separates feature participation, semantic contribution, and contextual refinement, allowing different sources to play different roles throughout the prediction process. EviSAM2 retains most of the pretrained SAM2 backbone and requires only 4.0M trainable parameters for multimodal adaptation. On MCubeS, the four-modality model achieves 55.46% mIoU, these results demonstrate the effectiveness of maintaining explicit source-wise evidence throughout multimodal foundation-model adaptation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.