acceptodds
Under review as a conference paper at ICLR 2027

AIRE: Closed-Loop Agentic Research for Text Embeddings

Abstract

Developing effective text embedding models extends far beyond executing a training script. It requires researchers to navigate interdependent decisions across data curation, objective design, sampling strategies, and evaluation protocols. Because key design trade-offs often emerge only after end-to-end training and retrieval evaluation, embedding development is fundamentally iterative. While recent advances demonstrate that full research workflows can be automated, doing so for text embeddings requires orchestrating domain-specific decisions across data pipeline construction, model training, and retrieval benchmarking. In enterprise settings, this automation must additionally handle distributed user roles, shared accelerator clusters, and strict data governance. To address these challenges, we introduce AIRE (Autonomous Investigation and Research for Embeddings), a multi-agent framework that structures these operations into an autonomous, persistent research loop. In AIRE, a central coordinator delegates literature review, data construction, implementation, training, and evaluation to specialized sub-agents, maintaining a shared cross-round state to guide iterative discovery. We also introduce AIRE-Bench, a benchmark for autonomous embedding research with three task environments built from public retrieval datasets. AIRE-Bench evaluates the resulting model and the research process, tracking model quality, improvement rate, exploration coverage, failure resilience, and compute cost. AIRE achieved an AIRE-Bench score of 89.61.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.