Agent Forest: Controlled Complexity Growth for RL Environments via a Multi-Agent Orchestration System
Abstract
Language models are increasingly expected to act as agents that use tools, follow user constraints, observe changing environment state, and complete long-horizon tasks. Training and evaluating these agents requires large collections of challenging, verifiable tasks, yet constructing such data at scale remains difficult. Existing synthetic task-generation methods frequently rely on single-pass prompting, producing tasks that are limited in complexity, weakly grounded in executable environments, and difficult to expand while preserving consistency. We introduce Agent Forest, a multi-agent framework for controlled complexity growth that systematically transforms trusted seed tasks into more complex descendants while preserving executability and verification guarantees. Agent Forest uses five specialized agents that increase complementary dimensions of interaction complexity: broader tool use, branching dependencies, long-range memory requirements, partial observability, and task composition. Rather than simply rewriting instructions to create superficial difficulty, these agents reason over available tools and environment state to modify the underlying interaction problem itself. Because transformations can accumulate and be applied in different orders, a single seed task can generate multiple descendants with distinct complexity profiles, forming a forest of task variants that grows both the scale and diversity of trusted agentic datasets. We apply Agent Forest to three agentic environments, -Retail, -Airline, and General-Agent, transforming trusted seed tasks into diverse descendants with varying complexity profiles. As interaction complexity accumulates, the resulting tasks become substantially more challenging for strong solver agents. On -Retail, Qwen 3.5 397B drops from 75.4% Pass@1 on the original seed tasks to 54.7% on fully transformed descendants. On -Airline, Gemma 4 31B drops from 66.0% to 28.6%, while GLM-5.2 drops from 72.0% to 53.6%. On General-Agent, Gemma 4 31B drops from 68.8% to 25.6%. Overall, accumulated complexity growth reduces Pass@1 by 18.4–43.2 percentage points relative to the original seed tasks, demonstrating that Agent Forest can systematically generate more complex agentic tasks while preserving executability and validity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.