acceptodds
Under review as a conference paper at ICLR 2027

Training Reasoning Models to Generate Diverse Scientific Ideas

Abstract

One of the first steps of a research pipeline is generating ideas, and good ideation means laying out several distinct directions. Large Language Models (LLMs) are increasingly used to automate this step, but most work optimizes the quality of individual ideas, overlooking diversity. In our work, we focus on increasing the diversity of LLM generated ideas, while making sure they are high quality. We study this for document-grounded ideation, where, given prior papers, a model must propose a set of ideas that are each high-quality and clearly distinct from one another. To optimize for this, we propose training with a set-level reinforcement learning objective. In a single generation, the model writes K structured idea cards grounded in the source papers, and the set is scored jointly, where each idea's quality is weighted by its distance from the ideas written before it. This teaches the model to steer each new idea away from its predecessors. Our trained 8B model achieves higher creative utility (a measure combining quality and diversity) than Claude Opus 4.8 and GPT-6 Astra on our contamination-free benchmark. In a pilot human study with four experts, our sets are judged as promising as Claude Opus 4.8's and contain marginally more unique ideas. When a coding agent runs proof-of-concept experiments on the ideas, our model's success@k is comparable to Claude Opus 4.8's, so we see no evidence that the added diversity costs executability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.