EvoPolicyGym: Benchmarking Coding Agents for Autonomous Policy Evolution
Abstract
Self-evolving agents that learn from interaction can improve their behavior by iteratively developing policies from environmental feedback. However, evaluating this capability remains challenging because agents may differ not only in the policies they produce, but also in how they explore the environment, allocate interaction budgets, and revise their policies over time. We introduce EvoPolicyGym, a benchmark framework for evaluating autonomous policy evolution under a standardized interaction budget. EvoPolicyGym assesses policy development trajectories and applies a validation protocol to select a policy from each run for independent evaluation on held-out episodes. The framework integrates 357 environment variants across 18 ecosystems. We evaluate four coding agents across all eight tasks spanning games and continuous-control domains. The results show that Astra achieves the highest mean final performance among the evaluated agents. Analysis of policy development trajectories further reveals that Astra iteratively identifies policy failures and incorporates the resulting evidence into subsequent policy revisions. These results establish EvoPolicyGym as a framework for studying how coding agents transform interaction experience into reusable executable policies, and provide evidence of emerging capabilities in autonomous policy development.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.