What Should the Missing Modality Say Structured Candidate Recovery for Incomplete Multimodal Learning
Abstract
Missing-modality recovery is often underdetermined: the same partial observa-tion can support several task-relevant interpretations of what an absent modal-ity would express. Deterministic reconstruction may collapse these alternatives,while stochastic recovery produces samples without explicitly associating themwith distinct semantic interpretations. We propose HyDiff, which represents thisambiguity with a finite, data-dependent set of structured candidates. A frozenmultimodal reasoner proposes candidates, a learned compatibility scorer assignsoperational weights, and candidate- and target-specific retrieval supplies instance-grounded priors. Conditional diffusion then recovers each missing target underthe retained candidates, after which task predictions are aggregated with theirweights. Candidate-weight entropy is used for selective prediction; the weightsare defined only over the reasoner-generated set rather than treated as a humaninterpretation posterior. The design therefore separates uncertainty over structuredinterpretations from the realization of each recovered modality and keeps compet-ing semantic alternatives explicit rather than leaving them as anonymous stochasticsamples. Experiments on three multimodal sentiment benchmarks evaluate feature-and modality-level missingness with matched-input, sampling, full-anchor, andreduced-anchor controls. When explicit target-facing fields are removed from theshared anchor, HyDiff retains a smaller but consistent advantage over matchedP-RMF, separating semantic assistance from the recovery architecture; additionaltransfer, completion, candidate-coverage, and efficiency diagnostics are reported inthe appendix.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.