Selecting Papers the Future Will Build On: Uptake-Weighted Coverage over Knowledge Threads
Abstract
Selecting a compact set of papers to understand a research field is a set-selection problem: a good reading list should cover distinct lines of work while prioritizing knowledge likely to remain relevant to future research. Existing approaches based on citations, graph centrality, or LLM importance scores largely evaluate papers independently, making them prone to redundancy and retrospective bias. We introduce WCov, an uptake-weighted coverage framework that decomposes papers into contribution-level claims and groups semantically related claims across papers into knowledge threads. We treat knowledge inheritance as the conceptual target and operationalize thread uptake as the number of later papers whose extracted claims semantically continue the thread. A citation-free thread-level predictor estimates uptake from semantic representations and thread metadata, after which weighted maximum coverage selects papers that cover high-uptake threads without double-counting shared content. The objective is monotone submodular, so greedy selection carries the standard approximation guarantee; an optional LLM-derived consensus prior trades future-uptake coverage against recovery of established landmark papers. We evaluate WCov on graph neural networks and federated learning, selecting from papers through , observing uptake in –, and holding out – for evaluation. The uptake predictor achieves AUCs of and , against and for extrapolating early uptake. At a budget of papers, prediction-weighted selection covers and of held-out semantic-continuation events, against at most and for any external baseline. Adding the consensus prior improves landmark recovery at the corresponding coverage cost. These results establish future-uptake-weighted coverage as a principled alternative to independently ranking papers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.