PreScience: A Dataset and Benchmark for Scientific Forecasting
Abstract
Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a renewable dataset and benchmark for scientific forecasting built around 147K recent research papers from multiple domains, together with companion papers covering author publication histories and citation links, yielding 1.06M papers in total. The resulting paper records include titles, abstracts, disambiguated author identities, influential references, topic labels, citation trajectories, and metadata snapshotted to respect temporal cutoffs. We instantiate seven exemplar tasks: five paper-anchored tasks—contribution generation, collaborator prediction, prior work selection, citation count prediction, and future combination prediction—and two aggregate topic trend forecasting variants. We develop baselines ranging from simple heuristics and embedding methods to frontier language models and agentic systems, and introduce LACER, an LLM-based metric for evaluating similarity of generated contribution descriptions that agrees better with human judgments than existing metrics. Finally, we compose task models into a 12-month synthetic corpus whose papers are systematically less diverse and less novel than human-authored research from the same period.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.