Reinforcing Agentic Creativity in Scientific Ideation with Night Science
Abstract
Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured _day science_ to loosely structured, serendipitous _night science_ that reaches ideas beyond those typically considered. We introduce **AI Night-Scientist**, an agentic framework that uses reinforcement learning to teach models _when_ and _how_ to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: _action_ (what to do and how creatively), _process_ (when to explore versus exploit), and _outcome_ (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying _what kind_ of creativity to pursue is critical. Overall, our results suggest that *creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs*.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.