Architectural Missing-Modality Robustness in Semantic Segmentation
Abstract
Multimodal semantic segmentation is typically evaluated with every modality present, but real-world deployment conditions often differ. A model trained on several modalities should therefore stay accurate on any non-empty subset of them at test-time. We propose Adapted Mixture of Modalities (A-MoM), which—instead of relying on auxiliary signals, training strategies, or model-specific losses—attains robustness architecturally through: (1) token-level, context-dependent mixing, since a sensor's advantage is usually local; (2) an initialization that keeps mixing low by default, so the model does not learn to depend on all modalities; and (3) low-rank adapters that translate features between modality pairs. Across 4 datasets spanning 8 modalities, A-MoM improves the average robustness score by 1.5 and retains on average 92.7% of the mIoU of instances of A-MoM specialized to each modality-subset on MCubeS, while also improving the average all-modalities score by 0.7. Thus, A-MoM narrows the gap from research to deployment under real-world conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.