MetaAgent-X : Breaking the Ceiling of Automatic Multi-Agent Systems via End-to-End Reinforcement Learning
Abstract
Automatic multi-agent systems aim to instantiate agent workflows without relying on manually designed or fixed orchestration. However, existing automatic MAS approaches remain only partially adaptive: they either perform training-free test-time search or optimize the meta-level designer while keeping downstream execution agents frozen, which creating a frozen-executor ceiling and leaving the end-to-end training of self-designing and self-executing agentic models unexplored. To address this, we introduce \name, an end-to-end reinforcement learning framework that jointly optimizes automatic MAS design and execution. \name enables script-based MAS generation, execution rollout collection, and credit assignment for both designer and executor trajectories. To support stable and scalable optimization, we propose Executor Designer Hierarchical Rollout and Stagewise Co-evolution to improve training stability and expose the dynamics of designer-executor co-evolution. \name consistently outperforms existing automatic MAS baselines, improving the six-benchmark average by 6.2 and 7.5 points over the strongest automatic MAS baseline on Qwen3-8B and Qwen3-4B respectively, and by 11.2 and 12.8 points over a single-agent baseline; it also outperforms single-agent best-of- and fixed multi-agent workflows evaluated under a comparable test-time token budget. Comprehensive ablations show that both designer and executor improve throughout training, and that effective automatic MAS learning follows a stagewise co-evolution process. These results establish end-to-end trainable automatic MAS as a practical paradigm for building self-designing and self-executing agentic models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.