acceptodds
Under review as a conference paper at ICLR 2027

Revela+: Adapting Multi-Stage Language Modeling Training for Reasoning Retrieval

Abstract

Reasoning-intensive retrieval is difficult because useful documents may support an answer through multi-step inference rather than resemble the query. Dense retrievers are mostly trained contrastively on annotated or synthetic query–document pairs, which reduces each document to a binary label on a similarity score: it records which document is relevant, but not how, or how much, it helps. A language model's next-token loss offers a finer signal, and Revela uses it to train a standalone dense retriever by letting the retriever weight the in-batch sequences that the model attends to. Revela, however, applies this only to pre-training, whereas language models are trained through a pipeline of pre-training, supervised fine-tuning (SFT), and preference tuning. Can a dense retriever be trained through this full pipeline, and what does each stage contribute? We introduce Revela+, which carries the retriever and the language model through all three stages: while attending to retriever-weighted sequences in its batch, the model predicts document text, then answers to queries, and then compares answers grounded in relevant and mined negative documents. No stage adds a document-ranking loss. Across five encoders from two model families, mean retrieval performance improves at both stage transitions on BRIGHT and BEIR. For Llama-3.1-8B, mean BRIGHT NDCG@10 rises from 25.1% after pre-training to 29.8% after SFT and 32.1% after preference tuning, surpassing contrastively trained ReasonIR-8B (24.4%) and RaDeR-7B (26.0%) with 6.6–7.7× less training compute than ReasonIR-8B.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.