acceptodds
Under review as a conference paper at ICLR 2027

ProgramGen: Scalable Environment Construction for Long-Horizon Agent Training

Abstract

Building an entire software repository from scratch is a challenging long-horizon task for software agents, requiring sustained exploration, implementation, execution, and debugging across many files. Training agents for such zero-to-one repository construction is essentially a data problem: realistic repositories are abundant, but scalable training tasks with reliable executable verification and high-quality long-horizon trajectories are not. We introduce ProgramGen, a repository-to-environment compiler for constructing such training data from existing software repositories. ProgramGen supports two complementary forms of reconstruction: ProgramGen-Doc retains the original tests as a hidden verifier and synthesizes an aligned specification, while ProgramGen-Exec retains the original program as an execute-only behavioral oracle and synthesizes behavioral tests from its observable behavior. Through agentic construction, ProgramGen produces 10,000 environments across five programming languages and 4,000 verifier-selected trajectories of up to 512K tokens. Training Qwen3.5-35B-A3B-Base on only 1,000 ProgramGen-Doc trajectories raises NL2Repo pass rate from 0.5% to 23.0%. The supervision also transfers beyond repository reconstruction and remains effective after post-training, improving Qwen3.6-35B-A3B from 28.8% to 31.8% on NL2Repo and from 69.6% to 75.4% on SWE-bench Verified.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.