Learning Role-Context Graph Policies for Adaptive Multi-Agent Collaboration
Abstract
Effective LLM-based multi-agent collaboration requires jointly deciding which agents should participate and how they should communicate. Existing methods often optimize communication over a prescribed team, potentially leading to redundant or insufficient participation for individual queries, or adapt participation through selection from available structures or sequential graph construction, limiting structural flexibility or decoding parallelism. Moreover, limited modeling of local directional role relationships can obscure the distinct effects of information reception and dissemination. To address these limitations, we propose RCGP, a query-conditioned Role-Context Graph Policy for adaptive multi-agent collaboration. To reflect their structural coupling, we formulate participation, functional roles, and directed communication within a unified graph-policy action space, allowing collaboration scale and topology to adapt jointly through parallel node–edge decisions. We further design directional role-context encodings to characterize local role diversity and neighborhood support separately in receiving and sending contexts, enriching graph-policy prediction beyond role identity and connectivity while preserving the functional asymmetry between information reception and contribution dissemination. We directly optimize the graph policy via execution-guided reinforcement learning (RL) with a structure-aware initialization. Together, the graph policy provides RL with a structured and relational collaboration space to optimize, while RL grounds the graph policy with task-aware execution utility, calibrating collaboration toward query-specific effectiveness. Extensive experiments on six benchmarks demonstrate that RCGP outperforms state-of-the-art baselines, while additional analyses show favorable cost–performance trade-offs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.