Privileged Posterior Calibration of Memory Routing for Incomplete Multimodal Learning
Abstract
Multimodal learning in real-world settings often faces incomplete inputs, requiring models to reason from only partially observed evidence. Recent prompt-based methods improve robustness by augmenting pretrained multimodal transformers with learnable prompts, while memory-based variants further retrieve instance-specific prompts to compensate for missing information. However, their retrieval decisions are typically conditioned only on the available modalities, making reliable memory access difficult when informative modalities are absent. We study the common setting in which complete multimodal pairs are available during training to simulate missing observations, while inference may receive only partial inputs. We propose Privileged Posterior Calibration (PPC), which uses the unmasked multimodal input as a training-only privileged reference for memory routing. PPC transfers complete-view routing distributions to an observed-view router through confidence-weighted dense calibration, while retaining sparse top-k retrieval for prompt injection and inference. Rather than requiring recovery of an unknowable complete-view route for every input, a conditional-projection analysis shows that the population objective learns the teacher routing predictable from observed evidence. Complete-view prototype initialization and semantic-neighborhood regularization further organize the shared prompt memory. Across MM-IMDb, Food101, and Hateful Memes under the RAGPT/ANGA-aligned 70% missing-rate protocol, PPC consistently outperforms our reproduced MemPrompt baseline in all nine settings. On image-missing MM-IMDb, PPC reaches 51.5 test F1-Macro, compared with 49.3 for complete-view feature distillation and 47.9 for an architecture-matched no-calibration control. These results show that complete-view supervision can improve memory-access decisions beyond alternative uses of privileged training information, while requiring only partial inputs and sparse memory retrieval at inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.