ResearchTree: Measuring the Research Horizon of Language Models
Abstract
How capable are language models at research? We measure one part of this question: how much of existing human research a model can reconstruct on its own, from a single step up to a whole paper. We introduce ResearchTree, which turns a research paper into a tree of self-contained problems, from its main result down to atomic steps with solutions taken from the paper. The model under test answers every problem independently, and a judge scores each answer against the paper's derivation beneath it. Two measurements follow. The research horizon is the problem size, in atomic steps, at which a model's expected score falls to one half. The decomposition curve shows how much of a paper a model recovers as the paper is given to it in smaller and smaller parts, and its deficit is the share the model still misses on average. On 20 recent theoretical papers and ten models, frontier models solve almost every atomic step but much less of each whole paper. Research horizons range from under one step to about 40, differ several-fold between fields, and across four leading releases have doubled roughly every eight months, a trend that rests heavily on the earliest release, while the decomposition deficit has halved about once a year. Posed a whole paper, models mostly leave parts of the derivation out rather than get them wrong. Because the trees are built automatically from new papers, the evaluation can be renewed after every model release and could extend to other fields. We release ResearchTree as a public dataset: 20 paper trees with 1,079 problems (744 atomic steps with reference solutions from the paper), 10,790 model answers and 32,880 graded judgments, together with the code to build new trees.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.