acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Structured Transformation: Mitigating Distribution Shift in Dense Retrieval Through Training-Time Preprocessing

Abstract

Dense retrieval models are trained assuming that finetuning on task-relevant queries improves performance, yet this assumption can break down when training data contains synthetic components or originates from a misaligned distribution with target tasks. We find that in such scenarios, naively finetuning on seemingly relevant data can result in negative transfer, causing significant degradations over not finetuning at all. We propose Adaptive Structured Transformation (AStrucT), an automatic preprocessing technique that leverages off-the-shelf Large Language Models (LLMs) to organize training documents into domain-specific structures prior to finetuning. These domain-specific schemas are induced automatically from a small sample of target-domain passages, without manual schema design or access to test-time queries. Across three model scales and twelve diverse domains (BRIGHT), AStrucT yields an average improvement of 3.75 percentage points (pp) nDCG@10 over direct finetuning, and 1.25 pp over the pretrained baseline, consistently mitigating negative transfer. For the 4B model, domain-matched templates outperform general rewriting and templates from other domains on average. These findings provide a practical strategy for adapting embedding models using automatically prepared training documents, without requiring LLM calls at inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.