acceptodds
Under review as a conference paper at ICLR 2027

When Does Graph-Based Retrieval Actually Help? A Benchmark for Scientific Literature Discovery

Abstract

Knowledge graph structure is widely expected to improve scientific literature retrieval, but existing evaluations either presuppose the graph or score a generated answer rather than the retrieved set, so the retrieval step itself is rarely the quantity measured. KGRAF-BENCH exposes one corpus through both a flat and a typed-graph view, retrieves with one encoder at one matched budget, and scores the retrieved paper sets against the citation decisions of survey authors. Its construction is automated: given a list of source-paper arXiv identifiers, the pipeline assembles a corpus, extracts a typed graph over it, and generates questions whose ground truth is the citations of the source papers. We release two ML/AI corpora of 8,219 and 89,476 papers with a graph over each, 647 questions, the per-question features, and a biomedical instantiation produced by the same pipeline. Our measurements identify the extraction schema as the component that determines which retrieval channel improves on flat retrieval. Two graphs over the same 8,219 papers at the same canonicalization threshold, differing only in what the extractor emits, place the advantage on different question sets, and the difference follows the paper-to-entity edges each schema provides. The advantage is consistent in direction across conditions and small in magnitude: over 26 configurations varying corpus scale, canonicalization threshold, encoder and construction method, entity-aware traversal is significantly ahead of flat retrieval in most and significantly behind in none. Selecting a strategy per question does not improve retrieval. A per-question oracle does not exceed its own permutation null, and a router trained on pre-retrieval features performs worse than the best fixed strategy. Tuning the extraction schema therefore matters more for a deployed system than increasing corpus size or adjusting the canonicalization threshold. We evaluate four retrieval strategies that isolate the contribution of entity embeddings, phrase extraction, and graph neighborhood expansion. The value of graph structure is conditional and predictable. Hybrid similarity over entity embeddings is the overall best strategy, winning on 80% of standard and 57% of relational questions. Graph neighborhood expansion is uniquely best on 24% of questions where strategies diverge. These are precisely the questions where the entity-paper gap, a computable pre-retrieval signal, is high. This signal distinguishes two retrieval regimes and predicts when graph structure adds value without any LLM calls. Per-question oracle selection yields R@50 gains of +0.09 over the best fixed strategy, demonstrating substantial headroom for learned adaptive routing. These findings are stable across 2K and 8K corpus scales. We release KGRAF-Bench to support research on retrieval for scientific discovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.