acceptodds
Under review as a conference paper at ICLR 2027

UNCERTAINTY-GATED KNOWLEDGE ALIGNMENT FOR ROBUST MULTIMODAL LEARNING UNDER MISSING MODALITIES

Abstract

Multimodal models degrade sharply when entire modalities are missing at test time. Existing remedies either reconstruct the missing stream, learn modality‑agnostic features, or supplement observations with external knowledge; however, knowledge is typically injected with fixed, availability‑count heuristics (e.g., w_K = 1/(1+|A|)) that ignore how unreliable the observed evidence actually is, and information‑bottleneck variants are often stated without a tractable derivation. We propose UGKA (Uncertainty‑Gated Knowledge Alignment), a framework that couples three components: (i) a modality‑symmetric variational information bottleneck (MVIB) with a closed‑form Gaussian ELBO that yields modality‑invariant latents; (ii) uncertainty‑weighted fusion (UWF), a precision‑weighted product‑of‑experts that provably minimizes the expected NLL among linear fusion rules under calibrated heteroscedastic heads; and (iii) an uncertainty‑driven progressive retrieval (UPR) expert that conditions the knowledge readout’s fusion precision on the predictive entropy of the fused observation, so knowledge competes with sensors exactly when the observed evidence is unreliable. We provide complete derivations and proofs, and we validate UGKA in a fully reproducible protocol on real sensor data (UCI HAR with three heterogeneous modalities) and on a controlled knowledge‑grounded benchmark (KGM3) with per‑sample reliability heterogeneity, reporting mean±std over three seeds, missing‑rate sweeps, per‑modality ablations, and confusion analyses. Our headline finding is quantitative: when retrieved knowledge is conditionally informative, fusing it as a full precision‑weighted expert cuts the full‑range degradation on KGM3 from 45.0 to 24.1 points (accuracy at r=0.8: 54.3 vs. 75.2), while on real sensor data— where no external knowledge source exists—modality‑dropout training remains the dominant factor, and our learned availability‑free weighting performs on par with, but not better than, the availability‑count heuristic. We report these positive and negative results side by side; every number is produced by the released code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.