Don't Drop the Weak Pairs: Alignement-Aware Distillation of Vision-Language Models
Abstract
Large-scale vision-language pretraining strongly benefits from data filtering, as poorly aligned image-text pairs provide noisy contrastive supervision and are often better discarded. We show that this intuition does not fully carry over to vision-language distillation. During distillation, weakly aligned pairs may indeed be harmful to the image-text contrastive objective but still provide valuable teacher supervision through feature, relation, or representation distillation. Rather than removing such samples, we propose to adapt how they are used. Specifically, we introduce an alignment-aware soft-gating mechanism that uses the teacher's image-text similarity to modulate, on a per-sample basis, alignment-sensitive objectives while preserving knowledge distillation. Well-aligned pairs receive stronger contrastive supervision, while weakly aligned pairs rely more heavily on the teacher signal. The gate is computed online from teacher embeddings already available for distillation, adding negligible overhead. Across CC12M and DataComp-M datasets, our approach consistently improves CLIP-KD and remains complementary to adaptive data selection with ACID. We further show that alignment weighting improves another distillation recipe, in TIPSv2, demonstrating that its benefits extend across different distillation settings. Our results suggest that, for vision-language distillation, noisy pairs are better reweighted than discarded. Code and pretrained weights will be open-sourced.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.