Task-Oriented Learning of Dynamic Communication Graphs for Multi-Round LLM Agents
Abstract
Multi-round LLM-based multi-agent systems solve tasks through repeated inter-agent communication, making who communicates with whom at each round consequential to the final answer. Most existing methods either fix the communication graph at inference or sample edges from behavior-independent distributions, preventing the topology from responding to evolving agent behavior. We introduce TodyComm, a task-oriented policy that constructs constrained, directed acyclic communication graphs conditioned on the unfolding interactions. The resulting graph policy is optimized end-to-end using the final task reward. We evaluate TodyComm in a dynamically adversarial setting, a controlled instance of shifting graph distributions, where agents may silently become unreliable at unknown rounds. Across six benchmarks, TodyComm outperforms fixed-graph and behavior-independent graph-learning baselines while retaining generalizability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.