acceptodds
Under review as a conference paper at ICLR 2027

What Fusion Accuracy Hides About Missing-Modality Robustness

Abstract

Multimodal models are usually evaluated with all modalities available, yet at deployment a modality is often missing. We show that fusion accuracy says little about this setting: models with nearly identical complete-input accuracy can differ sharply once one modality is removed, because complete-input training constrains only the joint output of the modality branches, not how each branch behaves alone. To locate such failures, we separate what the retained representation supports from what the model's own classifier extracts from it, using probes on frozen features and an independently trained unimodal model as reference. On audio–video, text–image, and clinical EHR–X-ray tasks, low missing-modality accuracy often hides a highly predictive representation; balanced training can even make the weaker modality's representation more predictive than unimodal training, while the classifier fails to use it. We propose Gated Residual Repair (GRR), a small affine correction applied only when a modality is missing, which recovers most of this gap without changing the encoders or complete-input predictions. Adapting the encoders helps only where the retained representation falls well below the unimodal reference, a case the diagnosis identifies in advance. Missing-modality robustness is thus a distinct property of multimodal models that fusion accuracy does not reveal.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.