acceptodds
Under review as a conference paper at ICLR 2027

ResearchMirror: Can AI scientists Match the Research Taste of Human Researchers?

Abstract

Research taste, the judgment of what to read, which problems to pursue, and which methods to try, increasingly shapes the work of AI scientists, yet it remains far less understood than their ability to execute research. We introduce ResearchMirror, a temporal benchmark for research taste that freezes the literature through 2025, generates research proposals from that snapshot, and compares them with work published by the same communities in 2026. Our experiments span six contemporary language models and four research areas from established to emerging fields, covering more than 90,000 proposals in total. AI scientists fall well short of human communities in both the literature they cite and the ideas they state, reaching about 26% of the human level on citation matching and only 18% on idea matching. Their populations can cite as diversely as a community and take up its tasks and methods, yet rarely arrive at the concrete approaches and research plans it pursues. Neither richer field context nor larger populations substantially close this gap. In blinded expert review, AI proposals read nearly as well as the 2026 papers yet are judged far less likely to succeed, and those whose ideas no human paper shares fare worst, suggesting weaker directions rather than overlooked opportunities. These results indicate that the main bottleneck in AI research taste lies less in finding the directions a community works on than in turning them into the concrete approaches it goes on to pursue. We will maintain ResearchMirror as an open, continuously updated leaderboard that benchmarks contemporary AI proposals against future publication cycles.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.