E-Orch: Towards Effective, Efficient, and Extensible Agentic Orchestration with Reinforcement Learning
Abstract
Agentic orchestration enables multiple autonomous agents to collaborate on complex tasks through adaptive task decomposition, delegation and execution. However, existing orchestrators largely rely on hand-crafted orchestration logic and prompting strategies, limiting their ability to adapt and generalise across tasks and executor configurations. In this work, we propose E-Orch, a reinforcement learning framework for learning effective, efficient, and extensible agentic orchestration through a workflow. Rather than planning all subtasks upfront, E-Orch organises execution around , a mid-level abstraction that scopes planning and execution around meaningful intermediate objectives, enabling the orchestrator to adapt its decisions as the task evolves. For each milestone, the orchestrator constructs a dependency-aware plan, delegates subtasks to suitable executors, and parallelises independent subtasks for efficient execution. We learn the orchestration policy from execution feedback, using milestone and plan decisions as natural units for fine-grained credit assignment. To estimate their downstream contributions, we construct tree-structured rollouts that compare alternative decisions under shared execution histories. The policy is jointly optimised for task performance and efficiency through complementary rewards for performance, execution cost, and planning completeness, including an uncertainty-aware performance reward that accounts for stochastic downstream outcomes. Across seven benchmarks spanning diverse domains, E-Orch achieves the best task performance under different agent configurations, improving on the strongest baselines by - points, and achieves - higher intelligence efficiency. The learned policy also transfers to new agent configurations introduced only at evaluation time, supporting the extensibility of agentic orchestration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.