AdaptWorld: Learning to Act and Adapt over Long-horizon Tasks in Dynamic Environments
Abstract
Long-horizon tasks in dynamic environments require agents to revise decisions as evidence arrives while managing earlier commitments. AdaptWorld constructs executable training environments that make this adaptation consequential. Public-information reference comparisons guide bounded recipe revision; independent assessment admits task families, and execution-and-outcome filtering selects model demonstrations. Across finance, grid storage, inventory, and network recovery, revision raises recipe acceptance from 34.4% to 72.7%. We fine-tune two Qwen backbones on 5,120 episodes from 4,226 tasks. Four-domain training improves both backbones on CEO, Merchant, and local Vending. For 27B, it exceeds the strongest single-domain source by 11.3%, 18.2%, and 69.4%, respectively, at a matched supervision target; 35B retains stronger task-specific specialization. Broader task coverage helps, while the best difficulty mixture depends on the target. At four times the reference horizon, SFT increases completion by 23–31 percentage points across all six model–benchmark pairs. These results link decision-oriented environment construction to transfer and sustained execution. Code is available at https://anonymous.4open.science/r/anonymous_code3-1025.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.