Dual-Guided Prompt Learning under Missing Modalities
Abstract
Existing prompt-based methods for incomplete multimodal inputs typically use a single fixed set of prompts across all missing patterns, without adapting them to the sample-specific task-relevant evidence provided by the available modalities. Moreover, they cannot effectively recover the complementary semantic information associated with the missing modality for each sample. To address these limitations, we propose Dual-Guided Prompt Learning (DGP), a framework that combines condition-guided prompt adaptation with prototype-guided semantic compensation. For condition guidance, we develop an evidence-aware hierarchical prompt tree, where the missing pattern determines the routing space and a sample-wise Fisher score dynamically selects and aggregates prompt experts. For prototype guidance, we build a learnable cross-modal prototype memory that uses the available modality representation to retrieve the corresponding prototype and inject it into the missing-modality proxy representation. Together, these two forms of guidance enable DGP to perform sample-wise prompt adaptation and compensate the missing branch with retrieved cross-modal semantic propotypes, without reconstructing the raw missing inputs. Experiments on three visual recognition benchmarks show that DGP achieves the best performance in 24 of 27 test configurations. Further evaluation on two multimodal sentiment analysis benchmarks demonstrates its broader applicability, yielding an average accuracy improvement of 2.03 points over the strongest baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.