acceptodds
Under review as a conference paper at ICLR 2027

ResearchTree: Measuring the Research Horizon of Language Models

Abstract

How capable are language models at research? We measure one part of this question: how much of existing human research a model can reconstruct on its own, from a single step up to a whole paper. We introduce ResearchTree, which turns a research paper into a tree of self-contained problems, from its main result down to atomic steps with solutions taken from the paper. The model under test answers every problem independently, and a judge scores each answer against the paper's derivation beneath it. Two measurements follow. The research horizon is the problem size, in atomic steps, at which a model's expected score falls to one half. The decomposition curve shows how much of a paper a model recovers as the paper is given to it in smaller and smaller parts, and its deficit is the share the model still misses on average. On 20 recent theoretical papers and ten models, frontier models solve almost every atomic step but much less of each whole paper. Research horizons range from under one step to about 40, differ several-fold between fields, and across four leading releases have doubled roughly every eight months, a trend that rests heavily on the earliest release, while the decomposition deficit has halved about once a year. Posed a whole paper, models mostly leave parts of the derivation out rather than get them wrong. Because the trees are built automatically from new papers, the evaluation can be renewed after every model release and could extend to other fields. We release ResearchTree as a public dataset: 20 paper trees with 1,079 problems (744 atomic steps with reference solutions from the paper), 10,790 model answers and 32,880 graded judgments, together with the code to build new trees.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.