Flow Matching for Soft-Label Debiasing in Long-Tailed Dataset Distillation
Abstract
Dataset distillation compresses large-scale datasets into compact synthetic data to reduce storage and training costs. In long-tailed scenarios, soft labels generated during distillation suffer from entangled bias originating from the imbalanced distillation model and synthetic images. Existing calibration methods typically assume this bias as a linear additive component and apply static statistical shifts to the logits. However, the distribution shift induced by class imbalance is highly nonlinear, rendering linear adjustments insufficient for disentangling semantic relationships from frequency-induced bias. This paper proposes Flow Matching Label Transport (FMLT), a framework that models soft-label debiasing as a continuous transport process on probability manifolds. We formulate the transition from biased soft-label distributions to unbiased target distributions using ordinary differential equations. By learning a time-dependent vector field conditioned on synthetic image features, FMLT transports the biased labels along nonlinear trajectories without relying on additive assumptions. Experiments on CIFAR-10/100-LT and ImageNet-1k-LT demonstrate that FMLT consistently improves performance across various distillation baselines and data budget settings. In extreme long-tailed scenarios, the nonlinear transport mechanism restores inter-class semantic topologies and improves tail-class accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.