acceptodds
Under review as a conference paper at ICLR 2027

POSEIDON: Scaling Training Environments for Scientific Agents from Code Repositories

Abstract

Training environments remain a major bottleneck for agents in data-driven scientific discovery. Existing collections of executable and verifiable training environments are small, and their construction remains difficult to scale. To address this limitation, we introduce POSEIDON, an automated pipeline for generating diverse, verifiable training environments from scientific code repositories at scale. From the same repository budget, POSEIDON extracts nearly 10× as many training environments as a prior approach by decomposing scientific workflows into four complementary task types: Reproduction, Subflow, Completion, and Correction, where each type targets a distinct capability for scientific research. A recursive improvement loop further refines failed candidates and recovers otherwise discarded environments. Using POSEIDON, we construct PEARL, a collection of 97,659 executable and verifiable environments from 7,110 public repositories spanning 10 scientific disciplines. To the best of our knowledge, PEARL is the largest and most broadly multidisciplinary collection of scientific-agent training environments to date. Qwen3.5 models (9B, 27B, and 35B-A3B) trained on PEARL achieve consistent gains across six challenging scientific-agent benchmarks. Overall, the scale and diversity of PEARL offer opportunities for training scientific discovery agents. Code is available at https://anonymous.4open.science/status/POSEIDON-89D7.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.