CodeEvoGym: Learning from Experience Across Software Releases
Abstract
Production web applications never stand still. Each release introduces features, fixes defects, and changes internal structure, making an agent's knowledge of the codebase increasingly outdated. This paper asks whether coding agents can continually update their understanding of a large web application as it evolves, and whether doing so improves their ability to implement changes in unseen releases. We introduce CodeEvoGym, a benchmark and training environment constructed from the release histories of six production open-source web applications spanning three programming languages. Each task provides the source tree immediately preceding a real release, along with its associated development record: release notes, pull requests, issues, and their discussions. The agent's patch is evaluated using the tests added or modified in the post-release source tree. During training, the agent may query a grading oracle that executes the gold test suite in a private container and returns the number and identities of failing tests, together with an LLM-generated summary of their error patterns. The agent distills the resulting feedback into three persistent artifacts, a memory store containing application-specific knowledge, a skill library containing reusable procedures and documentation that is readable by both humans and LLM agents. These artifacts are frozen after the training releases and evaluated on subsequent unseen releases. We examine the impact of using these artifacts on agent performance and cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.