Transporting Layerwise Structure for Text Embedding Distillation
Abstract
Large text embedding models achieve strong performance but incur high deployment costs, motivating knowledge distillation (KD) into lightweight models. Existing text embedding distillation methods primarily focus on matching final embeddings, leaving the student's intermediate computation underconstrained. We first observe that strong text embedding teachers share a consistent layerwise trajectory: pairwise distances expand into a pronounced hump in the middle layers before contracting, while the positive–negative margin widens rapidly in the final third of the network. However, output-only KD does not transfer this intermediate trajectory, and existing layerwise objectives also fail to recover it when applied to text embeddings. These observations expose three challenges: (i) failure to preserve relations among input texts, (ii) inability to align teacher and student networks with different numbers of layers, and (iii) neglect of the late-layer separation trend between positive and negative samples. To address these challenges, we propose Layerwise Structure Transport (LST), which (i) encodes relations among input texts in a common space, (ii) learns a soft correspondence to align teacher and student networks of different depths, and (iii) transfers ranking preferences with greater weight at depths where positive–negative separation emerges. Experiments show that LST reliably transfers the teacher's layerwise trajectory and consistently improves student performance under all tested Output KD objectives. Our code can be found in the Reproducibility Statement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.