CollabOrchestra: Learning to Orchestrate Human-Agent Teams
Abstract
AI agents are increasingly deployed in occupational contexts, where many real-world scenarios involve humans and agents working together. Effective human-agent collaboration can outperform either party alone while preserving human agency, but requires understanding each other's capabilities and deciding when humans should delegate, verify, or intervene. We present **CollabOrchestra**, a practical method for learning an orchestrator policy that delegates work across a human-agent team. CollabOrchestra casts orchestration as *adaptive workflow plan generation*, where the orchestrator constructs and revises a sequence of steps specifying goals, assignees, and handoff deliverables. Our training proceeds in two stages: we first profile agent workers through sandboxed execution and learn a feasibility predictor from their execution traces, then train the orchestrator through multi-turn reinforcement learning to jointly optimize plan quality and execution feasibility. Across GDPval, JobBench and their underspecified variants, CollabOrchestra outperforms prompting-based orchestrators in both plan quality and execution outcome. The resulting team also outperforms the best solo agent, improving Task Performance from 0.474 to 0.744 on GDPval and from 0.260 to 0.286 on JobBench. Our analysis shows that CollabOrchestra learns to assign humans diverse, task-dependent roles, including directly performing work when appropriate. In a user study on participants' own tasks, users prefer CollabOrchestra-trained plans 16% more and report high collaboration legibility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.