IdeaReward: Benchmarking and Training Reward Models for Research Idea Generation
Abstract
Research idea generation aims to produce research ideas that can be developed into publication-worthy research projects. However, as large language model (LLM)-based systems generate research ideas at scale, the central bottleneck shifts from producing candidate ideas to reliably evaluating, selecting, and improving them before costly research execution. Reward models offer a promising solution for this goal by learning to assign quality-aware reward scores from preference data. In this paper, we present a systematic study of reward modeling for research idea generation. We present IdeaRewardBench, the first multi-signal, living benchmark for research idea evaluation. IdeaRewardBench contains 3,185 idea preference instances grounded in complementary signals of idea quality, and is updated semiannually to reflect the latest research frontier. We evaluate a comprehensive set of 23 baselines on IdeaRewardBench, and find that all baselines struggle to distinguish high-quality research ideas: the best baseline achieves only 62.8% pairwise accuracy on the review-grounded split and 60.9% on the impact-grounded split, where 50% is the random baseline. We then present IdeaReward, the first reward model for research idea generation, trained on over 20,000 high-quality idea preference pairs constructed by a fully automated pipeline. IdeaReward establishes state-of-the-art performance on IdeaRewardBench, outperforming baselines based on much larger LLMs. We further show that IdeaReward improves downstream test-time best-of-N idea selection. We will release our data, models, and code upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.