acceptodds
Under review as a conference paper at ICLR 2027

Cross-Modal Correspondence Matters: Unified Prototype-Guided Data Synthesis for Multimodal Dataset Distillation

Abstract

Multimodal dataset distillation aims to construct compact yet informative image–text datasets that preserve essential knowledge from large-scale datasets for efficient cross-modal learning. Recent prototype-guided approaches typically construct visual and textual prototypes separately, primarily preserving modality-specific structures while insufficiently capturing the joint dependency that defines image–text correspondence. In this paper, we analyze this limitation from an information-theoretic perspective and introduce UPSyn, a Unified Prototype-guided Synthesis framework that designed to adequately preserve cross-modal information from the original dataset. Specifically, we construct two cross-modal common spaces from correlation and geometry perspectives, and jointly cluster the paired data by integrating the structural evidence from both spaces, thereby producing informative prototypes that better maintain image–text dependency. Additionally, we leverage a large language model to aggregate and summarize the entities surrounding each textual prototype into a refined prototype, providing semantically rich guidance for subsequent distilled data generation. Extensive experiments on Flickr30K and MS-COCO demonstrate that our method consistently outperforms coreset-based, pixel-optimization-based and generative-based baselines in bidirectional retrieval, while retaining strong cross-architecture generalization with an efficient training-free paradigm.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.