acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Modality Fusion at Test Time: Label-Free Selection for Robust Multimodal Sensing

Abstract

Multimodal sensing benefits from complementary observations, but distribution shift can change the utility of individual modalities, making indiscriminate fusion detrimental. Selecting useful modalities without target labels is challenging: incorrect predictions can remain confident, and cross-modal agreement can be inflated by shared predicted-class preferences. We propose Test-Time Modality Selection (TMS), a label-free framework that selects a shared modality subset for an unlabeled target batch. Using cached predictions from classifiers frozen after source training or few-shot adaptation, TMS combines prediction coverage screening, cross-modal support estimation through empirical-marginal-corrected agreement, and reference-based marginal contribution assessment. Auxiliary calibration supports coalition comparisons, while final predictions use unweighted averaging of the selected modalities' original logits. Selection requires no source calibration data, target labels, parameter updates, or additional classifier forward passes. Experiments on MM-Fi and XRF55 evaluate cross-environment transfer under zero-, one-, and two-shot adaptation and controlled modality corruption. TMS improves mean zero-shot accuracy on MM-Fi from 32.16% to 58.79% and one-shot accuracy under severe mmWave corruption on XRF55 from 20.60% to 35.12%. These results demonstrate the value of reassessing modality participation at test time using unlabeled target predictions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.