acceptodds
Under review as a conference paper at ICLR 2027

Resolving Semantic Degeneracy in Multimodal Dataset Distillation

Abstract

Diffusion-based methods for dataset distillation (DD) cluster images into prototypes and decode them with diffusion models into a synthetic set, becoming a prominent DD paradigm due to their training-free, architecture-agnostic design. Existing methods transfer this pipeline to multimodal dataset distillation (MDD) on paired image–text data by clustering each modality separately at the target budget and keep the DD decode process unchanged. However, per-modality clustering ignores cross-modal interactions and collapses distinct pairs, budget-sized clustering merges fine-grained semantics into meaningless prototypes, and standard diffusion guidance renders samples from distinct prototypes indistinguishable. We term these failures semantic degeneracy and propose DPRG (Distinction-Preserving Representation and Guidance). On the representation side, we assign each pair jointly and over-cluster before selecting back to the budget, keeping semantically distinct pairs in separate prototypes. On the guidance side, we denoise each target against its hardest image- and text-side alternatives, keeping distinct prototypes apart during synthesis. Extensive experiments on Flickr30K and MS-COCO demonstrate that DPRG outperforms existing baselines across all evaluated distillation budgets and downstream architectures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.