Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
Abstract
Dense-caption retrieval requires distinguishing images paired with highly similar, detail-rich descriptions. We empirically examine contrastive fine-tuning from a strong pretrained dual encoder and observe that the standard InfoNCE loss can become small on sampled batches while held-out retrieval errors remain. Through an analysis of gradient dynamics, we find that InfoNCE rapidly loses optimization signal: within the first epoch, its loss falls below on of sampled batches, and over the run its gradient norm below in of measurements, even as held-out retrieval continues to improve. We relate this premature saturation to the abundance of near-duplicate captions in dense-caption datasets, where the most confusable negatives remain unresolved. We therefore introduce HN-CLIP, which uses the text encoder's own text–text geometry to construct per-negative adaptive similarity margins. A detached caption-similarity matrix is added to the negative logits, assigning larger margins to more similar captions without mining, synthesizing, or resampling negatives. The objective requires only one similarity matrix and a masked logit addition during training, with no auxiliary data, additional parameters, offline preprocessing, or inference-time overhead. Across four dense-caption retrieval benchmarks, HN-CLIP improves over the strongest competitors by – Recall@1 while training faster than GOAL and faster than StructXLIP. It also improves all six tested fine-tuning frameworks in-domain and surpasses full-data GOAL and StructXLIP with only of the DOCCI and DCI training data, demonstrating strong data efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.