Building Trustworthy Scientific Environments for Coding Agents
Abstract
Automated scientific discovery is emerging as a new paradigm where coding agents interact with scientific environments to pursue open-ended research goals. While recent work has improved agent capabilities, the role of environments in shaping scientific exploration remains largely unexplored. In this work, we show that environmental feedback influences how coding agents explore and optimize scientific objectives through controlled studies with four coding agents across four feedback settings. We then introduce SciEnBuilder, an automated framework that transforms paper–repository pairs into trusted scientific environments by connecting scientific claims, executable evidence, and protected evaluation. Using SciEnBuilder, we construct 30 scientific environments for coding agents from a large collection of top-tier papers. Human evaluation confirms their faithfulness to the original scientific goals and the reliability of their evaluation. Compared with raw repositories, these environments improve research efficiency, eliminate observed shortcut behaviors, and enable stronger long-horizon optimization. These results highlight the importance of trustworthy scientific environments for reliable automated scientific discovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.