Ballpark Mining: Region-Based Contrastive Distillation for Embeddings
Abstract
When compressing large embedding models through knowledge distillation, aligning student and teacher representations alone can leave substantial gaps in downstream performance. Contrastive objectives over hard negatives mined from the teacher’s embedding space are therefore commonly added to supply additional supervision. However, hard-negative mining suffers from a well-known limitation: the presence of false negatives. In particular, in duplicate-rich corpora, the nearest candidates to an anchor are often its paraphrases, so the contrastive loss pushes apart texts with the same meaning. Existing remedies can require similarity labels or auxiliary models. To address this, we propose Ballpark Mining, a label-free contrastive distillation method that uses teacher-derived regions to select negatives. Using the variability of a frozen teacher’s embeddings under stochastic perturbation, we define for each anchor a region of representations that the teacher cannot reliably distinguish from it. Candidates within this region are treated as probable false negatives, and the nearest candidates outside it are selected as negatives. We evaluate our proposed method with teachers ranging from 110M to 27B parameters on a deduplicated Wikipedia corpus and a duplicate-rich Quora corpus, using semantic similarity and retrieval tasks from MTEB and BEIR. On Quora, distilling a 110M-parameter teacher into an 11M-parameter student with Ballpark Mining reduces the fraction of duplicate negatives from 21.9% to 2.3% and increases duplicate recall@10 from 0.919 to 0.973. The gains extend to a 27B teacher, where Ballpark Mining improves the average score across semantic similarity and retrieval tasks from 0.211 with standard hard-negative mining to 0.285 when distilled on Quora. On Wikipedia, where duplicates are rare, performance remains comparable to standard hard-negative mining. In the 110M-to-11M setting, estimating regions at an intermediate teacher layer further improves the average score by approximately 3% on Quora and 6% on Wikipedia relative to final-layer regions. Together, our studies show that the perturbation behavior of the teacher provides a simple, label-free criterion for negative selection in embedding distillation, while identifying a corpus-transferable choice of region size as a key challenge.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.