acceptodds
Under review as a conference paper at ICLR 2027

Revisiting Vision-Language Embedding Distillation Through Negative Composition

Abstract

Contrastive learning is a core objective for vision–language embedding models, where in-batch examples define the negative relations that shape the learned space. In heterogeneous training, globally mixed batches introduce negatives from datasets whose candidate pools are evaluated separately. We study this negative-source mismatch in vision–language embedding distillation. Across FastVLM-0.5B and LLaVA-OneVision-0.5B, classification and visual question answering, five objectives, and three training seeds, task-consistent batches improve overall Precision@1 by 1.54–3.72 points relative to the corresponding mixed-task objective. For contrastive-only training, the gains are 2.03–3.04 points. By comparison, under task-consistent batching, auxiliary distillation changes the overall score by only to points relative to contrastive-only training. Dataset-level effects are heterogeneous, with large gains on many datasets and regressions on a few. Because the samplers handle per-dataset remainders differently, the comparison does not isolate the causal effect of negative-source composition. Nevertheless, the aggregate association is consistent across settings, identifying negative-source composition as a first-order experimental axis and motivating matched-exposure and gradient-level tests of the proposed mechanism.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.