acceptodds
Under review as a conference paper at ICLR 2027

StrataDistill: Structure-Aware Distillation for Long-Document Retrieval

Abstract

Long-document retrieval often depends on evidence confined to a small part of a document. Single-vector embedding models encode documents without the query, making local relevance difficult to preserve amid unrelated content. We introduce StrataDistill, a distillation method that uses document structure to guide supervised fine-tuning (SFT) of embedding models. Frozen teachers combine full-document similarity with the strongest local match to construct relevance distributions over candidate documents. A shared student learns these distributions through full-document scoring alongside supervised contrastive learning. Using source-specific teachers consolidates supervision from separately adapted models in the same retriever. The method reuses existing retrieval data and document structure without query-to-section annotations, and also supports short documents without decomposition. Across five long-document datasets and four embedding backbones, StrataDistill with source-specific teachers improves macro-average nDCG@10 by 2.85 to 4.66 points over mixed-task SFT. Further analysis shows a smaller gap between full-document and local-view scoring, consistent with partial transfer of local relevance into document representations. Inference preserves the original index size and scoring procedure, storing one vector per document.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.