acceptodds
Under review as a conference paper at ICLR 2027

A Study of Downstream Rewards for Retrieval-only Adaptation

Abstract

Reward design is a central challenge in learning multi-step retrieval policies from downstream feedback. We investigate how different reward signals enable retrieval-only adaptation, where a compact embedding model learns to select evidence while the answer-generating language model remains frozen. Using Q-RAG as a common framework, we compare rewards based on generated answers, reference likelihoods, and semantic uncertainty, with supporting-fact supervision as a reference. We also introduce Group Information Gain, a contrastive likelihood reward over answer candidates evaluated offline. Experiments on HotpotQA and 2WikiMultiHopQA characterize how these signals affect evidence retrieval and answer quality. Using the selected reward, we then train the retrieval policy over a Wikipedia corpus containing approximately 21 million passages and evaluate it on seven QA datasets. Results show that retriever-only fine-tuning improves downstream answer quality and achieves competitive performance with fine-tuned LLM-based search policies while updating only a compact embedding model. Adapting the retriever already required by the RAG pipeline thus offers a flexible way to specialize search while leaving the LLM unchanged for other tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.