acceptodds
Under review as a conference paper at ICLR 2027

Loop Reinforcement Learning for Scientific Discovery

Abstract

Large language models (LLMs) have demonstrated strong search capabilities when given well-defined problems and quantitative verifiers. Advances in both model capabilities and orchestration harnesses have accelerated progress in scientific discovery. However, LLMs tend to generate similar candidates, concentrating exploration within narrow regions of the search space and risking premature convergence that limits the breadth of discovery. Recent approaches seek to broaden exploration by orchestrating LLMs with evolutionary algorithms, using test-time scaling to maintain diversity and expand search coverage. Despite their effectiveness, these approaches rely on manually designed search rules, heuristics, or agentic workflows that may constrain exploration. To address this limitation, we propose Loop Reinforcement Learning (Loop RL) for scientific discovery. We first introduce a simple loop harness consisting only of sequential iteration, parallel execution, and fully asynchronous shared memory for exchanging information across parallel trajectories. Beyond these basic mechanisms, the harness leaves exploration strategies to the model, providing a minimal baseline with ample flexibility to explore and learn. We then develop Loop RL to optimize the discovery performance of this loop through credit assignment across iterations and parallel trajectories. Across scientific discovery tasks, Loop RL achieves faster search speed and higher-quality solutions, attaining state-of-the-art results on several tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.