EnvOpt: Closed-Loop Environment Optimization for Agent Post-Training
Abstract
Data-centric recursive self-improvement (RSI) seeks to improve a model by repeatedly turning its evolving behavior into new training data. For tool-using agents, sustaining this process is more difficult because the data are executable environments: each round must preserve a coherent and tight connection from observed behavior to training intent, executable realization, and realized learning outcomes. Our analysis finds that even a capable agent frequently loses this connection, showing an incomplete understanding of model behavior, unfaithful implementation of intended environments, and ineffective use of feedback across iterations. To address these challenges, we introduce , a framework that organizes environment evolution around traceable training targets. It combines a traceable behavioral map to expose recurring patterns, iterative environment construction to faithfully realize training targets as executable environments, and target-conditioned context refinement to facilitate accumulated feedback utilization. This turns environment generation from a sequence of isolated attempts into a closed-loop data improvement process that adapts as the target model changes. Extensive experiments on -Bench, BFCL-v4, ACEBench-Agent, and AppWorld show that yields substantial improvements in the target model's performance, while the vanilla agent baseline shows no consistent improvement trend.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.