acceptodds
Under review as a conference paper at ICLR 2027

ReOrg: Post-Orchestration for Adaptive Multi-Agent Systems via Graph-Level Relative Advantage Policy Optimization

Abstract

Large language model agents are increasingly deployed as collaborative teams, and performance depends on the organization that channels their interaction as much as on the capability of individual members. Existing multi-agent systems follow a pre-orchestration paradigm: a collaboration graph is designed once, from the input task alone, and executed unchanged. This design-then-execute assumption overlooks a fundamental asymmetry: many coordination failures, such as missing capabilities, correlated errors across members, and deadlocked review loops, become observable only during execution, precisely when runtime feedback can diagnose them. We propose post-orchestration, a paradigm that treats this runtime evidence as a first-class organizational signal: the system executes the initial graph, scores the resulting draft, revises the collaboration topology, and re-executes. We instantiate post-orchestration as ReOrg, a learned reorganization layer that stacks on any pre-orchestrator and edits the collaboration graph through a small set of task-agnostic structural actions; task-specific repair is expressed through the role prompts of newly added nodes rather than hard-coded actions, and a selective trigger restricts intervention to error states, skipping the large majority of already-correct drafts. The reorganization policy is trained with RO-GRPO, an extension of group-relative policy optimization to graph-level organizational decisions. Across eight benchmarks, ReOrg stacked on the strongest recent pre-orchestrators yields non-negative gains on all 24 method×benchmark combinations, with seven reaching statistical significance under paired testing and gains of up to 4.6 points; a policy trained on code generation alone transfers to reasoning benchmarks and to unseen pre-orchestration architectures, improving an AgentVerse pipeline by 8.5 points—evidence that what is learned is the skill of adjusting the organization itself, not a fixed organizational form. Gains peak at moderate baseline accuracy and shrink toward zero at saturation, positioning post-orchestration as error recovery rather than capability enhancement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.