acceptodds
Under review as a conference paper at ICLR 2027

MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning

Abstract

Collaboration topology determines the performance and execution cost of multi agent systems based on large language models. Existing representative topology generation methods typically adopt a single collaboration granularity, using either individual agents throughout the topology for fine grained control or predefined groups to reuse established collaboration patterns. However, even within the same task, different functional roles may require substantially different levels of collaboration. A fixed collaboration granularity therefore fails to capture the diverse collaboration demands across functional roles. Our key insight is to flexibly select the collaboration granularity for each functional role, thereby combining the fine grained control of individual agents with the reusable collaboration patterns of predefined groups. Introducing mixed granularity, however, exponentially enlarges the search space for topology generation. Existing topology generation methods typically rely only on final task outcomes as reinforcement learning rewards. Such sparse rewards provide insufficient intermediate feedback in the enlarged mixed granularity search space, which can lead to unstable convergence or slow topology generation. To address this challenge, we propose MAGIC, a dense reward reinforcement learning framework for mixed granularity graph generation. MAGIC constructs a mixed granularity MAS topology incrementally. For each functional role, it determines whether the role should be instantiated as an individual agent or a reusable group, and then connects the instantiated unit to the existing MAS topology. We directly optimize the construction policy using returns from trajectories sampled by the current policy. We further introduce a reward shaping mechanism based on structural complexity, structural redundancy, and utility. This mechanism provides intermediate feedback through sampled utility signals and structural signals while preserving the cumulative task reward. Experimental results show that MAGIC achieves the highest mean score among the evaluated methods across eight benchmarks and achieves strong inference efficiency in the efficiency evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.