acceptodds
Under review as a conference paper at ICLR 2027

Light-DiffRetriever: Budget-Flexible Retrieval with Diffusion Language Models via Matryoshka Representative Tokens

Abstract

Diffusion language models can produce multiple retrieval representations in one forward pass, but *fixed-budget training does not supervise effectiveness under truncation*. We introduce **Light-DiffRetriever**, a budget-flexible extension of DiffRetriever that **learns Matryoshka representative tokens: ordered vector prefixes supporting different query and passage budgets from one checkpoint**. Learnable slot offsets with a decaying-scale initialization provide an ordering bias, while a curriculum gradually introduces prefix-level contrastive supervision alongside full-budget retrieval training. The resulting representations can be truncated after encoding, reducing dense-index storage and late-interaction computation without retraining or re-encoding the corpus. On Dream and LLaDA-1.5, reducing the passage budget from 16 to 4 vectors retains 98.2% and 98.0% of full-budget MS MARCO MRR@10, and 98.4% and 97.8% of BEIR-7 nDCG@10, respectively, with one quarter of the dense-vector storage. These results support nested representative-token learning for resource-adaptive retrieval.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.