Neighbor Co-Occurrence Alone Should Not Define Supervision in Embedding Distillation
Abstract
Relational knowledge distillation for text embeddings is typically performed within mini-batches, making batch composition determine which teacher-defined relationships the student can observe. Random batches may miss a seed text's close teacher neighbors. Grouping these neighbors makes the relevant comparisons available, but grouping alone does not tell the student which neighbor the teacher considers closer to the seed text, or by how much. We introduce Neighborhood Similarity Transfer Knowledge Distillation (NeST-KD), which turns each mini-batch into a shared teacher-neighborhood subgraph. NeST-KD samples seed texts, retrieves their nearest neighbors in the teacher embedding space, and then matches teacher and student conditional similarity distributions over the resulting subgraph. This transfers graded differences in similarity while allowing other eligible texts to provide additional supervision from the same encoded examples. An auxiliary objective also matches teacher and student similarity distributions from each seed text to the other examples in the shared subgraph. Extensive experiments on nine standard sentence-embedding datasets show that NeST-KD improves over the strongest baseline by and absolute points for the 66M and 22M MiniLM students, respectively. Controlled studies show that correctly assigning teacher similarities, beyond merely selecting neighbors, improves downstream performance and relational fidelity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.