Multi-Agent World: Scaling Multi-Agent Orchestration Training via Recursive Environment Improvement
Abstract
Multi-agent systems assign complementary subtasks to different agents, and their effectiveness depends on an orchestrator that partitions a request into subtasks, schedules them against their dependencies, and integrates the evidence that workers return. Training such an orchestrator requires multi-agent environments, in which a task provides subtasks that can be distributed across agents and a verifier that reports how they were distributed. Existing agentic environments are synthesized for a single agent, and they yield mostly linear call chains under outcome-only verification. An orchestrator trained inside them therefore has no scheduling decision to learn and no feedback on the decomposition it chose. We present Multi-Agent World (MAW), an end-to-end framework that develops multi-agent orchestration by coupling multi-agent environment synthesis with multi-agent reinforcement learning. MAW introduces recursive environment improvement, which recursively composes reference task graphs through serial and parallel operators inside single-agent environments, and turns each composed graph into one coherent request through semantic reconstruction and execution-grounded admission. The same composition record identifies which subtasks are verified as independent, which is what a schedule-aware process reward needs in order to score the trajectory of the whole team. Trained under these process rewards, the orchestrator learns to decompose a request, run independent work concurrently, and integrate what the workers return. Experiments on five agentic benchmarks show that our trained orchestrator, MAW-8B, improves its backbone by 8.9 points on average. Further ablations on the challenging DeepPlanning benchmark show that MAW-8B can surpass Claude Sonnet 4.6 as an orchestrator.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.