Does Agentic RAG over Knowledge Graphs Help Scientific QA? Only Where No Single Query Reaches the Evidence
Abstract
Retrieval-augmented generation grounds a language model in a corpus, but one retrieval call answers one kind of question. Questions that bridge two papers, aggregate a quantity reported across several, or combine a numeric detail with an entity relationship need evidence a single dense query does not return. Graph-based retrieval already offers the operations such questions require: dense vector search, entity-centred graph traversal, and relationship-centred synthesis. A deployed system, however, fixes one of them per query. We present ModeRAG, in which a Planner assigns each sub-task the retrieval mode its part requires and declares the dependencies between sub-tasks, an Executor runs ready sub-tasks in parallel, and a Reporter answers from an evidence chain that records the chunks, graph paths and entities behind each step. We release LHCpubs-100QA, 100 multi-paper questions over LHC papers with annotated evidence, and evaluate alongside AI4EIC2023, 50 mostly single-paper questions. ModeRAG improves answer correctness by 7.7 points over the same pipeline restricted to vector search, with the largest gains on questions that bridge or aggregate across papers; the difference is unresolved on the single-paper benchmark. Dense retrieval remains the backbone, and decomposition costs 24,268 tokens per question.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.