DACE: Divide-and-Compose Evidence for Missing-Modality Recognition
Abstract
Multimodal recognition must remain robust when modalities are unavailable at test time. Existing methods either reconstruct entire missing representations or fuse modalities without distinguishing transferable from modality-specific information. We introduce Divide-and-Compose Evidence (DACE), a lightweight recognition head that achieves robustness through learned evidence separation and selective cross-modal transport. For each observed modality, a complementary router divides its encoded feature into two branches: specific evidence that captures modality-unique cues and remains local, and shared evidence that expresses transferable cross-modal information. When a modality is missing, DACE recovers only its shared slot by transforming and averaging the shared branches from observed sources, leaving modality-specific content unestimated. The recovered and observed evidence is then composed and refined through availability-conditioned attention and observed-pair agreement. Experimental results demonstrate consistent robustness: across three benchmarks and 27 missingness conditions, DACE achieves the highest score in every comparison, and a single trained model maintains strong performance when test rates shift from 10% to 90%. On naturally incomplete medical data, DACE reaches 84.30% accuracy and 75.60% AUROC, demonstrating robustness on real-world incomplete medical data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.