acceptodds
Under review as a conference paper at ICLR 2027

When Multimodality Stops Helping: Auditing Representation-Conditioned Residual Modality Utility

Abstract

Whether an auxiliary modality helps a prediction is not a property of the modality: it is residual to what the model’s primary representation has already learned. A strong primary representation can leave a powerful auxiliary modality with nothing to add, whereas cross-modal reconstructability tempts exactly the wrong inference. We formalise this view with two estimands. Incremental modality gain (IMG) is the actual test-risk reduction from adding the auxiliary modality under a matched low-capacity head family, with a separate parameter-count control; residual modality utility (RMU) asks whether the component of the auxiliary signal that cannot be predicted from the primary representation explains the target error the primary model left. We show that cross-modal predictability implies neither conditional sufficiency nor gain. Across representation-level ladders spanning weak/generic to more task-relevant encodings on a clinical glaucoma structure–function task, CMU-MOSI, and AV-MNIST, IMG concentrates at weaker or less adapted levels and collapses toward zero at stronger, more adapted levels, while RMU tracks realised gain on every domain (seed-cluster bootstrap Spearman , 95% CI [0.483, 0.825] pooled; per-domain CIs exclude zero on both public benchmarks) and survives negative, missingness, and capacity controls. An independent sweep over seven clinical backbones shows the same pattern. RMU is a capacity-conditional diagnostic, not conditional mutual information. These results support representation-conditioned residual utility as an operational diagnostic, not a universal sufficiency test.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.