Topic Is Not Agenda: What Text Embeddings Miss in Scientific Retrieval and What Citation Structure Adds
Abstract
Research agents retrieve candidates by text similarity and read only the top of the list, so what text similarity misses, they never see. In this paper we measure how far text similarity is from finding papers that share a study's problem and approach, using the documents that authors cite as gold, and we propose a method that uses the citations of nearby papers to close part of this gap. We test on a public must-cite benchmark (MasterSet), a 4M-paper corpus in eight scientific domains, and US patents. The top 100 by text similarity miss 61% of the most essential papers an author cites, the baselines, benchmarks and methods its experiments build on. In the 4M-paper corpus, 30.8% of the cited papers that it holds lie beyond cosine rank 1,000, where no reranker of the text candidates can reach them, and updating the query vector does not help on average. Text embeddings find the topic of a paper, not its specific problem and approach. The method, neighbourhood citation retrieval, asks the query's nearest text neighbours what they cite, adds those documents as candidates, and ranks them by the neighbours' citations. Its training-free form, the neighbour vote, cuts these misses to 43%, and a learned form cuts them to 35%. On the same MasterSet candidates the vote beats citation popularity and a trained Citeomatic-style reranker, and adding the vote and related neighbourhood signals to that reranker raises R@100 from .556 to .660. Fine-tuning the encoder on citations does not replace the vote, which still adds to every fine-tuned encoder we tested. We also analyse where the gain comes from with several controls, including the removal of the query's own organisation (its co-authors or its assignee) from its neighbours. Because it needs no training and uses new citations as soon as they appear, the approach can serve as a retrieval interface for research agents in collections where cosine similarity is saturated, such as scientific papers and patents, where many documents share a topic and text alone cannot tell which of them share a study's problem and approach.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.