acceptodds
Under review as a conference paper at ICLR 2027

Breaking the Orchestration Monoculture: Training Generalizable LLM Agents via Multi-Orchestration Reinforcement Learning

Abstract

Reinforcement learning (RL) has become the dominant paradigm for training LLM-based agents on tool-use tasks, yet virtually all existing work trains under a single, fixed orchestration such as REACT throughout the entire RL process. This orchestration monoculture limits both generalization and data utilization efficiency: agents overfit to one reasoning pattern and become brittle on out-of-distribution queries that demand different problem-decomposition strategies. We argue that the root cause runs deeper than a static choice: the training value of an orchestration drifts as the policy's capability grows, so early leaders may cap the long-run ceiling or even collapse while complex scaffolds only pay off once the policy can exploit them. We therefore treat orchestration as a self-evolving curriculum: a fixed pool of structurally diverse orchestrations whose training-time allocation is routed online by the policy's own reward stream. Instantiated as Multi-Orchestration Reinforcement Learning (MORL) with Discounted-UCB routing, the curriculum reweights rather than rewrites, adapting to capability drift while keeping the rollout distribution stable. On seven Search QA benchmarks, MORL outperforms a rollout-matched single-orchestration baseline by 5.3 exact-match points on average (43.3% vs. 38.0% with Qwen2.5-3B-Ins), with notable gains on OOD multi-hop sets (+7.8 on Bamboogle, +5.7 on MuSiQue), demonstrating strong task generalizability and robust orchestration adaptability across model sizes and domains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.