StrataDistill: Structure-Aware Distillation for Long-Document Retrieval
Abstract
Long-document retrieval often depends on evidence confined to a small part of a document. Single-vector embedding models encode documents without the query, making local relevance difficult to preserve amid unrelated content. We introduce StrataDistill, a distillation method that uses document structure to guide supervised fine-tuning (SFT) of embedding models. Frozen teachers combine full-document similarity with the strongest local match to construct relevance distributions over candidate documents. A shared student learns these distributions through full-document scoring alongside supervised contrastive learning. Using source-specific teachers consolidates supervision from separately adapted models in the same retriever. The method reuses existing retrieval data and document structure without query-to-section annotations, and also supports short documents without decomposition. Across five long-document datasets and four embedding backbones, StrataDistill with source-specific teachers improves macro-average nDCG@10 by 2.85 to 4.66 points over mixed-task SFT. Further analysis shows a smaller gap between full-document and local-view scoring, consistent with partial transfer of local relevance into document representations. Inference preserves the original index size and scoring procedure, storing one vector per document.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.