Ideas Have Genomes: A Comprehensive Benchmark for Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
Abstract
Auto-research systems can retrieve papers, propose hypotheses, and draft research text, but current evaluations rarely ask whether a system understands how scientific ideas descend from prior work. We introduce (Genome-based Evolutionary liNeage Evaluation), a benchmark for evaluating scientific lineage reasoning and lineage-grounded idea generation. The benchmark is built on an explicit representation: a gene is a fine-grained, typed idea atom such as a mechanism, problem niche, limitation, observation, or claim; an is the typed bundle of genes carried by a paper or proposal; and a GenomeDiff aligns genes across predecessor-successor pairs. contains 1,961 golden lineage traces, 1,085 paper genomes, and 920 pairwise GenomeDiffs. It instantiates two complementary evaluations: (42 task types, 1,029 instances) tests closed-form lineage reasoning across genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification; tests whether generated proposals can be inserted as coherent descendants under Question, Library, and Lineage information settings. Experiments on 15 LLM-based scientists show that current systems struggle with compositional lineage reasoning (GPT-5.5 reaches 23.1% exact accuracy; the best CLI harness reaches 27.3%) and that structured lineage context changes which systems are competitive in generation rather than uniformly improving all systems. reframes evaluation of automated research around what can be claimed from a proposal's relation to prior scientific lineages.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.