Diffusion Pretraining for Data-Efficient Text Embeddings
Abstract
We compare bidirectional masked language modeling (MLM), bidirectional absorbing-mask discrete diffusion, and causal next-token prediction using parameter-matched ModernBERT-style models at four scales from 20M to 500M parameters, each pretrained on the same Ettin data mixture with a nominal budget of approximately 20 input tokens per parameter. We evaluate unsupervised SimCSE, frozen-backbone adaptation, and supervised InfoNCE fine-tuning on 10k, 100k, or 1.06M AllNLI and MS MARCO triplets, alongside raw-representation and random-initialization controls. We evaluate on seven STS benchmarks, four BEIR datasets, 20-Newsgroups clustering, typo robustness, and all 41 tasks of MTEB English v2. Under full supervision, the objectives reach similar MTEB means; the causal model leads local STS and clustering at every tested scale, and local retrieval from 60M onward. The objectives nevertheless separate in two ways. First, the diffusion–MLM retrieval gap narrows at larger scales and reverses at 500M on BEIR retrieval. Second, at 60M–500M, both bidirectional models outperform the causal baseline on local retrieval under SimCSE or 10k-triplet fine-tuning. Diffusion leads at 500M, with BEIR nDCG@10 of 0.229 versus 0.217 for MLM and 0.128 for the causal model under SimCSE, and 0.250 versus 0.243 and 0.135, respectively, with 10k triplets. With frozen-backbone adaptation, diffusion leads local STS and retrieval at 60M and 150M, whereas the causal model leads at 500M. Our results suggest that objectives that tie on aggregate embedding quality can differ systematically in retrieval, and that the direction of the difference depends on scale and on how much supervision the adaptation stage provides.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.