ResponsiveEnv: Evolving Training Worlds for LLM Agents through Online Editing
Abstract
LLM agents learn by interacting with their environments. Recent work on environment scaling seeks to improve agent learning by expanding the number and diversity of training environments. We instead obtain more useful experience from existing worlds by revising their conditions while preserving task intent. A limited pool of original worlds can then provide varied practice for the same tasks. Useful conditions depend on the learner's current capabilities. We therefore propose ResponsiveEnv, a framework that revises world conditions throughout training based on current learner trajectories. An environment editor simplifies or hardens selected worlds, and validated revisions replace them for subsequent training. We evaluate ResponsiveEnv on unedited ALFWorld and τ²-bench tasks. On ALFWorld, we edit 100 original worlds with Claude Opus 4.8, GPT-5.6 Sol, or a distilled 8B editor. Each editor raises the unseen-task success rate by 18 to 19 percentage points over static training on the same worlds. Editing these 100 worlds also outperforms training on 2,288 static worlds under the same learner-update budget. For the 8B editor, replacing the current learner's trajectories with success rates alone, or with the initial policy's trajectories, lowers the unseen-task success rate by 6.3 and 5.5 points, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.