acceptodds
Under review as a conference paper at ICLR 2027

Do LLMs Have Research Taste? An Analysis Using BóLèBench

Abstract

Large language models (LLMs) are increasingly used for recursive self-improvement, in which they generate and implement research ideas intended to advance their own capabilities. With finite compute, researchers must decide which ideas to pursue before their outcomes are known. This judgment reflects research taste developed through experience. Researchers increasingly rely on LLMs to make these decisions, yet their research taste remains poorly measured. We introduce BoleBench, a benchmark that grades foresight with hindsight. We start from parallel rollouts, where different agents attempt to improve a shared baseline on a fixed set of ML research tasks under matched compute budgets. An LLM reads abridged, outcome-free descriptions of two attempts at the same task and predicts which scored higher. The executed outcomes serve as ground truth without requiring human outcome labels. BoleBench is a pipeline that turns executed attempts into test items covering five categories of research decisions, tests pick judgments at four information levels, and scores each judge on multiple axes. New item banks are built from post-cutoff corpora as models retrain. On AI4AI pairs with a clear boldness gap, a simple pick-the-bolder heuristic outperforms every full-description judge. On the subset where the bolder attempt loses, no judge exceeds chance. We release BoleBench, 4,059 blinded pairs from four corpora. On the ten shared anchor pairs, human and model majorities made identical choices, including the same three errors. On the 16-judge leaderboard the newest frontier model leads at 70.8%, against a nine-judge majority of 67.1%. On AI4AI, judges add little accuracy beyond the boldness heuristic. Asked to pick the best of ten attempts at one task, the top judge is right 36.0% of the time against a chance rate of 10%, so current judges have much room to improve.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.