acceptodds
Under review as a conference paper at ICLR 2027

FORAGE: Related Works Prediction as a Benchmark for Agentic Retrieval

Abstract

Agentic Retrieval Augmented Generation, where language models autonomously plan queries against a retriever, has several important applications, but lacks rigorous evaluation: single-turn benchmarks ignore iteration, multi-turn benchmarks reward fact lookup over document understanding, and web-based benchmarks are prone to contamination. We introduce FORAGE (**F**inding **O**ut **R**elated **A**rticles via **G**enerative **E**xploration), a live agentic retrieval benchmark built on a related works prediction (RWP) task: given a research paper with its related works section, citations, and bibliography redacted, the system must query a scientific corpus to rank candidate papers by likelihood of being cited as related work. We instantiate FORAGE separately on NeurIPS and ICLR 2025 papers and evaluate three open-source LLMs along with a leading close-source model across sparse, dense, and late-interaction retrievers. The best configuration recovers nearly a third of each paper's cited related works, and an oracle probe that removes the retriever from the loop recovers substantially more, placing the bottleneck in query generation and retrieval. FORAGE provides substantial headroom for future work on agentic retrieval, and is designed to be refreshed each conference cycle to remain contamination-resistant as models evolve. We release the dataset [here](https://anonymous-hf.com/a/o5o42hifzs1a/).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.