acceptodds
Under review as a conference paper at ICLR 2027

ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments

Abstract

Academic paper search is a core step in scientific research, and LLM-based search agents are emerging as a promising paradigm for iterative, intent-driven literature exploration. However, existing benchmarks are insufficient for systematically evaluating agentic academic search under realistic open literature environments. We propose ScholarQuest, a taxonomy-guided benchmark for agentic academic paper search. ScholarQuest is constructed over 1,000 computer science topics and organized around four research intents, including method-oriented, setting-anchored, comparison-based, and scope-controlled queries. It further provides scalable answer construction and a shared retrieval backend ScholarBase for reproducible evaluation. We evaluate at both the system and model levels, with nine representative literature search systems and six leading foundation models, respectively. The latter are evaluated using a shared multi-turn search harness. Benchmarking results show that agentic methods outperform single-shot retrieval baselines, yet the best-performing agent among these systems achieves only 0.312 Recall@100 and 0.352 Recall@All, indicating substantial room for improvement. In addition, analyses of search efficiency, intent-level robustness, and failure cases further highlight the benchmark’s ability to provide multi-dimensional evaluation signals for academic paper search agents. Our code and data are publicly available。

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.