CM: Cross-Modal Mode Learning for Generative Multimodal Dataset Distillation
Abstract
Multimodal dataset distillation (MDD) compresses a large image-text corpus into a small synthetic set on which a vision-language model can be trained. Generative MDD methods cluster each modality separately and then pair image and text clusters by a post-hoc assignment; nothing in two independent clusterings forces cluster to denote the same concept in both modalities, and the non-differentiable matching cannot correct this. We present CM, which makes cross-modal correspondence part of mode discovery rather than a step after it. CM learns image and text modes, one pair per slot of the distillation budget, in a frozen vision-language embedding space. Its core is cross-modal swapped prediction: the image of a pair must predict the mode assignment of its caption and vice versa, with assignments balanced over the slots, so that index comes to denote one concept in both modalities. A mode-level contrastive term aligns same-index modes and keeps the two banks distinct, and a mode-anchoring term ties each mode to the raw-scale centroid of the samples it owns, so the learned modes are decoded directly and are themselves the distilled set. CM requires no offline clustering, no post-hoc matching and no generator fine-tuning. It raises cross-modal index agreement by 30-85% over post-hoc matching and reduces mismatched image-caption pairs from 7.2% to 0.5% on Flickr30K at the largest budget. Across two benchmarks, four budgets and two evaluation backbones unseen during distillation, it improves mean retrieval recall over the strongest generative MDD baseline in all 16 settings, by +3.6 on Flickr30K and +2.61 on MS-COCO on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.