Agents Don't Need a New Index: A Scientific Knowledge Layer over Lexical Search
Abstract
The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, and nowhere more so than in science, where agentic pipelines are proliferating. Large scientific corpora are already reachable through mature lexical-based search engines, but these interfaces are not built for AI agents: they expect keywords and complex syntax and return whole papers, so every agent must learn the query syntax, issue several searches, and read entire articles to find the evidence it needs. We introduce LIBRARIAN, a knowledge layer that upgrades lexical-based search interfaces for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the search engine, then reads the selected papers and locates the relevant evidence. We instantiate LIBRARIAN on Europe PMC, an open lexical-based search engine that provides comprehensive access to scientific literature, and we evaluate multiple AI agents with access to LIBRARIAN across four scientific benchmarks: literature synthesis, claim verification, open-form question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, agents equipped with LIBRARIAN improve Citation F1 by more than points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus by points, and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about points higher when using the LIBRARIAN rather than web search. Our results show that existing knowledge infrastructure can be reused in AI pipelines without reindexing corpora or sacrificing performance. We release LIBRARIAN openly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.