acceptodds
Under review as a conference paper at ICLR 2027

Measuring Cross-Condition Generalization Under Missing Modalities: A Closed-Form Bound and a Cautionary Case Study

Abstract

Missing-modality robustness methods are typically evaluated at a single, matched train/test missing rate - a protocol that cannot distinguish a model that has genuinely learned to generalize across missingness patterns from one that has merely overfit to the pattern it was shown. We derive a closed-form, optimal-transport-solver-free upper bound on the risk gap between any two missing-rate conditions, exploiting a structural fact specific to the missing-modality setting: the label is invariant under masking, so no irreducible “distortion” term is required, unlike general conditional-distribution-shift bounds for generative models. We verify the transport inequality numerically on synthetic data and use its closed-form geometry to define descriptive sensitivity diagnostics on real trained multimodal networks. These diagnostics support a lightweight, pre-registered evaluation protocol for cross-condition robustness claims. Applying this protocol to a recent missing-modality defense (BALM) across two backbones and two datasets, and to a second, structurally different method (AMB-DSGDN), we find a heterogeneous pattern that no single evaluation setup reveals: (a) a real, previously undocumented artifact in BALM's evaluation masking (a positional correlation with the fixed test-set ordering); (b) on one backbone, BALM's central sensitivity-reduction claim appears confirmed at a small sample size but does not survive a properly powered, pre-registered multi-seed evaluation, and a second dataset shows the same null pattern; (c) on a second backbone, a distinct, slope-independent absolute-risk-reduction effect looks large and unanimous at 3 seeds, then at the full pre-registered 6-seed batch is significant by paired -test but not by Wilcoxon - the same failure mode caught twice on two different axes; and (d) a structurally unrelated method (adaptive modality dropout) improves both slope and intercept together. No accuracy-at-a-few-rates table, the field's current norm, can distinguish these effects, and no single backbone/dataset/method combination reveals the full picture. We argue this decomposition, and the multi-seed, multi-axis protocol needed to see it, is a general need for the field, and offer both as a contribution independent of any specific verdict on any one method.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.