acceptodds
Under review as a conference paper at ICLR 2027

It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning

Abstract

Recent work increasingly trains large language models (LLMs) for retrieval. Yet existing methods predominantly optimize generation solely on the query side, leaving item representations fixed. This one-sided adaptation restricts mutual alignment and caps overall retrieval performance. We introduce , a co-evolving reinforcement learning framework that jointly trains both query-side and item-side LLM generators. Each generator produces a compact set of discrete keywords for retrieval, preserving full compatibility with existing keyword-based infrastructure. first employs supervised fine-tuning to establish a warm-started, aligned keyword space. It then applies co-evolving reinforcement learning, alternately optimizing each generator against the opposite side’s frozen index. Both sides optimize the same query-to-item objective: the query side receives directly, while the item side receives a counterfactual marginal reward designed to isolate each item's impact on retrieval performance. Evaluated against 10 representative sparse, dense, and generative retrieval baselines, achieves the strongest overall retrieval performance on both an internal APP Marketplace dataset and the public WANDS benchmark, improving over the strongest baseline by and , respectively. Further analysis demonstrates stable co-evolution dynamics and the emergence of more specific, balanced keyword representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.