From Capability to Utility: Reliability Configuration Calibration for Multimodal Fusion
Abstract
Existing research on modality imbalance mainly aims to improve the learning and classification capabilities of individual modalities. However, stronger unimodal capability does not directly ensure that correct evidence is used in fusion. We identify and characterize the capability-to-utility gap (C2U gap): when one modality predicts correctly and the other incorrectly, the fusion model may still make an incorrect prediction. By examining the joint correctness patterns of unimodal predictions, we find a mismatch in reliability configurations between in-sample and held-out data. In particular, critical configurations in which the dominant modality is wrong and the other modality is correct are underrepresented in in-sample predictions. Based on this diagnosis, we propose Reliability Configuration Calibration (RCC). RCC uses configuration deficits between in-sample and held-out calibration data to determine the directions and probabilities of feature replacement. Reliability-weighted counterfactual feature replacement provides additional fusion supervision for critical cases, while unimodal anchoring and adoption preservation help maintain unimodal capability and limit interference with previously reliable configurations. Extensive experiments show improved fusion accuracy across six benchmarks and higher visual conditional utility on three audio–visual datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.