acceptodds
Under review as a conference paper at ICLR 2027

Retrieval-Augmented Label-Attentive Prompt Learning for Multimodal Classification with Missing Modalities

Abstract

Missing modalities make multimodal classifiers brittle because the available evidence is partial, while common remedies either synthesize the missing modality or rely on static missing-aware prompts. We argue that missing-modality classification should instead be treated as retrieval-augmented evidence selection: a memory bank can provide useful neighbors, but the retrieved labels and embeddings are noisy and should not be trusted uniformly. We introduce Retrieval-Augmented label-attentive prompt Learning for Multimodal Classification with Missing Modalities (RALAP), a frozen-LLM framework that converts retrieved multimodal neighbors into dynamic soft prompts. RALAP first encodes the target self-embeddings and target–neighbor difference embeddings through a sparse autoencoder bottleneck to generate differential soft prompts. It then computes a label-attentive prompt by weighting retrieved instances with a similarity kernel and aggregating their labels into a target conditioned label prior. The resulting self, difference, and label prompts are injected into a frozen Qwen2-VL model, which jointly generates a class label and a short explanation. We further show theoretically that, with sufficient retrieval coverage and stronger aggregated evidence for ground-truth labels than distractors, ground-truth labels receive higher attention scores. Empirically, the results support the need for target conditioned reweighting and are consistent with the theoretical behavior of label attention. RALAP-VLM achieves the best reported performance in 10 of 12 evaluation settings across MM-IMDb, HateMemes, and Food101, with a maximum observed improvement of 21.84 percentage points on MM-IMDb. These results support the effectiveness of combining retrieval-enhanced sparse differential prompts with target conditioned label-attentive prompts under modality missingness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.